Skip to content

feat: endpoint describe panel (#97) - #189

Open
MathiasVDA wants to merge 4 commits into
mainfrom
claude/happy-rubin-fh9b41
Open

MathiasVDA wants to merge 4 commits into
mainfrom
claude/happy-rubin-fh9b41

Conversation

@MathiasVDA

@MathiasVDA MathiasVDA commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #97

Summary

Adds an Endpoint overview panel for getting to know an unfamiliar SPARQL endpoint while writing queries. Open it with the new Describe endpoint button in the tab control bar (also in the mobile menu) or with F8.

Panel behaviour

  • Docked to the right of the editor.
  • Can be pinned (stays open, restored after a reload), collapsed to a thin bar, or resized by dragging its left edge.
  • An unpinned panel collapses when the user clicks into the editor.
  • On small screens it opens as a bottom sheet.
  • Follows the endpoint of the active tab.

Service Description and VoID

  • Fetched automatically when the panel opens:
    • the SD by requesting the endpoint without ?query;
    • VoID from /.well-known/void, or embedded in the SD.
  • Shows features, extension functions, named graphs, dataset statistics, vocabularies and class/property partitions.
  • When an endpoint publishes neither (e.g. GraphDB), the panel says so instead of showing an error.
  • Descriptions are read up to 5 MB and parsed up to the last complete statement; the panel says when a description was cut off. Wikidata's SD is about 19 MB.

Overview queries

  • 20 queries across all six categories from the issue comment:
    • overview (named graphs, triples per graph, triple count, namespaces)
    • schema (classes, properties, SKOS schemes, SHACL/ShEx shapes)
    • instances (samples per class, hubs)
    • labels and languages
    • links (referenced hosts, linking properties, linked entities)
    • time and geo (temporal and geo properties, CRS, lat/long bounding box, and a bounding box computed from a sample of WKT literals)
  • Queries run only when clicked ("Run" or "Run all").
  • Every query has a LIMIT.
  • Lists such as named graphs are paginated with "Load more", so an endpoint with millions of graphs is not listed in one request.
  • Full-scan queries are labelled "may be slow".
  • At most two describe queries run at once, each with a timeout (60 s by default).
  • Queries run through the tab's request config and authentication, but without the tab's default/named graphs.
  • Incomplete results are marked Partial result. Virtuoso "anytime queries" (DBpedia) answer 206 / X-SQL-State: S1TAT, often with an empty result.
  • Error messages are made readable:
    • HTML error pages are reduced to their title;
    • rate limiting (429) is explained, including Retry-After;
    • Java stack traces (Blazegraph/Wikidata) are reduced to their root cause.

Remembered results

  • Results are cached per endpoint and survive other queries, tab switches and reloads. Each shows when it was fetched.
  • The cache has its own storage key, capped at 20 endpoints and 200 rows per result. MatGUI's quota handler would otherwise wipe the whole configuration if the cache filled the storage.

Working with results

  • Clicking an IRI inserts it into the query as a prefixed name and adds the missing PREFIX line.
  • Any describe query can be opened in a new tab.

Configuration

  • New endpointDescribe option: enabled, queries, categories, timeoutMs, maxConcurrentQueries, fetchMetadata.
  • Queries and categories can be extended, e.g. with endpoint-specific queries.

API additions

  • Tab.runBackgroundQuery() and Tab.getRequestInit().
  • yasgui.endpointDescribe.
  • A skipGraphArgs option on yasqe's executeQuery.

Other

  • n3 added to the yasgui dependencies (already used by yasr).
  • Changeset: minor.

Documentation

  • User guide: new "Endpoint Overview" section (including partial results, rate limits and large descriptions), and an F8 entry under keyboard shortcuts.
  • Developer guide: "Endpoint Describe Configuration" section with an example, plus API reference entries for yasgui.endpointDescribe, tab.runBackgroundQuery() and tab.getRequestInit().
  • README and website intro: feature bullet and config option.

Testing

Live endpoints. Every query was run against DBpedia (Virtuoso), Wikidata (Blazegraph) and data.europa.eu (Virtuoso). SD/VoID discovery was run against all three, and the panel itself was run in headless Chromium against Wikidata and data.europa.eu.

  • All queries are accepted by all three endpoints, except named-graphs / graph-sizes on Wikidata, which runs in triples mode.
  • Full-scan queries are refused or time out on the large endpoints, as their "may be slow" label says.
  • Fixed based on these runs:
    • the namespaces regex matched the empty string, which Virtuoso rejects;
    • partial 206 results looked like complete results;
    • the 19 MB Wikidata SD was downloaded in full;
    • blank-node datasets were shown as IRIs;
    • HTML and stack-trace error bodies were shown raw;
    • SKOS schemes was not flagged as slow;
    • runQuery() did nothing on a closed panel.
  • The ERA endpoint (data-interop.era.europa.eu) could not be reached from the development environment.

Unit tests:

  • query catalogue: bounded queries, balanced syntax, declared prefixes, no REPLACE patterns that match the empty string, LIMIT/OFFSET pagination, per-class sample generation, WKT bounding box;
  • SD/VoID extraction: unavailable/HTML/404 cases, de-duplication, blank nodes, size cap and truncation;
  • response handling: partial results, HTML/429/504/Blazegraph errors;
  • cache: persistence, truncation, eviction, storage failures;
  • skipGraphArgs.

Browser tests (Puppeteer, against a mocked endpoint):

  • SD/VoID display;
  • results kept after a regular query;
  • IRI insertion with its prefix;
  • following the active tab's endpoint;
  • pinned panel restored after a reload;
  • partial result and 429 messages;
  • auto-collapse when the editor gets focus.

Other checks:

  • Every generated query parses with sparqljs (checked ad hoc, not added as a dependency).
  • Local run on the branch with main merged in: npm run build, npm run unit-test (180 passing) and npm run puppeteer-test (47 passing).

🤖 Generated with Claude Code

https://claude.ai/code/session_01DEbQuDzCkbQKKsARAXTpKu

Add an "Endpoint overview" panel that helps to understand a SPARQL
endpoint while writing queries. It is opened with the new "Describe
endpoint" button in the control bar (or F8) and docks next to the
editor. It can be pinned, collapsed to a thin rail and resized, and on
small screens it opens as a bottom sheet.

- Reads the SPARQL Service Description (endpoint without ?query) and
  VoID (/.well-known/void or embedded in the SD) and shows features,
  extension functions, named graphs, dataset statistics and
  class/property partitions. Missing sources are reported, not errors.
- Bounded overview queries for all categories of the issue: named
  graphs (paginated), triple counts, namespaces, classes, properties,
  SKOS schemes, SHACL/ShEx shapes, sample instances, hubs, languages,
  links and time/geo (incl. a bounding box computed from WKT samples).
  Queries only run on demand; expensive ones are flagged.
- Results are cached per endpoint in their own storage key, so they
  survive new queries, tab switches and reloads, without risking the
  main configuration when the storage quota is exceeded.
- Clicking an IRI inserts it (prefixed, adding the PREFIX) into the
  query; describe queries can be opened in a new tab.
- Queries and categories are configurable via `endpointDescribe`.
- Tab gets `runBackgroundQuery` / `getRequestInit`, and yasqe's
  `executeQuery` a `skipGraphArgs` option.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEbQuDzCkbQKKsARAXTpKu
…eference

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEbQuDzCkbQKKsARAXTpKu
Findings from running the panel against DBpedia, Wikidata and
data.europa.eu:

- The namespaces query used a REPLACE() pattern that matches the empty
  string, which Virtuoso rejects. Use "[^#/]+$" and add a regression test.
- Virtuoso "anytime queries" (DBpedia) answer HTTP 206 / X-SQL-State
  S1TAT with incomplete, often empty results. Such results are now marked
  as partial instead of looking like "no data".
- Error pages are summarized: HTML bodies (Wikimedia 429, gateway 504)
  are reduced to their title, 429 explains the rate limiting (with
  Retry-After), and Java stack traces (Blazegraph) are reduced to their
  root cause.
- Wikidata's service description is ~19 MB. SD/VoID documents are now
  read up to 5 MB, parsed up to the last complete statement and flagged
  as truncated; partition de-duplication no longer is quadratic.
- Blank-node VoID datasets are no longer shown as IRIs and datasets
  without any information are dropped.
- SKOS concept schemes is flagged as possibly slow (timed out on
  data.europa.eu).
- runQuery() now uses the endpoint of the active tab even when the
  panel is closed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEbQuDzCkbQKKsARAXTpKu

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add describe endpoint features

2 participants