Skip to content

feat(yasqe): support SPARQL 1.2 syntax - #188

Merged
MathiasVDA merged 4 commits into
mainfrom
feat/sparql-1.2
Sep 29, 2026
Merged

MathiasVDA merged 4 commits into
mainfrom
feat/sparql-1.2

Conversation

@MathiasVDA

@MathiasVDA MathiasVDA commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Fixes #2

Adds SPARQL 1.2 support to the Yasqe query editor, so SPARQL 1.2 queries are highlighted and syntax checked instead of being flagged as errors. The grammar follows the SPARQL 1.2 Query Language Working Draft of 21 September 2026.

What's supported now

  • VERSION "1.2" declarations
  • Triple terms <<( s p o )>> in triple patterns, VALUES and expressions
  • Reified triples << s p o ~ r >>, reifiers (~) and annotation blocks {| ... |}
  • Language tags with a base direction, such as "hello"@en--ltr
  • The new functions LANGDIR, STRLANGDIR, hasLANG, hasLANGDIR, isTRIPLE, TRIPLE, SUBJECT, PREDICATE and OBJECT
  • ! applied to any unary expression (!!BOUND(?x)), and aggregates wherever a built-in call is allowed (e.g. ORDER BY COUNT(?x))

The spec's grammar notes only allow a reifier or annotation after a simple predicate (an IRI, a or a variable). The editor now reports this as an error when a property path such as :p/:q comes first, with an explanatory tooltip.

Commits

  1. Fix the grammar generator for current SWI-Prolog. Since SWI-Prolog 8.3, ==> and => are built-in rule operators. The generator therefore produced an empty table, so the grammar could not be changed. Productions are now stored as ebnf/2 / bnf/2 facts. Regenerating the existing SPARQL 1.1 grammar with the fixed generator gives a byte-identical _tokenizer-table.js.
  2. SPARQL 1.2 grammar and tokenizer support, with tests. Two existing bugs are also fixed:
    • [ :p ?o ] :a/:b ?c was reported as a syntax error. It is valid SPARQL 1.1.
    • The tokenizer's state.complete flag was always false. Nothing in Yasqe read it, but the new tests use it.
  3. Version-neutral file names: sparql11-grammar.pl → sparql-grammar.pl and gen_sparql11.pl → gen_sparql.pl. The public "sparql11" mode name is unchanged.

Testing

  • New unit test suite test/unit/yasqe-sparql12-grammar-test.ts with 82 cases:
    • valid SPARQL 1.2 queries
    • SPARQL 1.1 regression queries
    • invalid SPARQL 1.2 queries, each of which must be rejected on the expected line
  • Run against the previous grammar table, 40 of these tests fail, so they exercise the new rules.
  • New browser tests in test/run.ts:
    • the editor marks a valid 1.2 query as valid and an invalid one as invalid
    • formatting a 1.2 query with either formatter keeps it valid
  • All 109 complete example queries from the SPARQL 1.2 spec pass. The spec's other examples are fragments, use undeclared prefixes (which Yasqe flags on purpose), or contain a typo.
  • npm run unit-test: 154 passing. npm run puppeteer-test: 40 passing.

Known incompatibility

A < comparison written without a space before a full IRI, such as FILTER(?o<<http://example.org/a>), used to be accepted and is now reported as a syntax error. SPARQL 1.2 reads << as a single token (the longest match wins), so a SPARQL 1.2 parser rejects this query as well. Adding a space fixes it: FILTER(?o < <http://example.org/a>). This is also noted in the changeset.

Notes for reviewers

  • The spec forbids blank-node syntax in DELETE DATA, DELETE WHERE and DELETE clauses. That is enforced for explicit forms (_:b, [], ~ _:r). A reified triple without an explicit reifier implicitly creates a blank node; it is still allowed there, because the spec text only covers explicit syntax.
  • Existing non-standard leniencies are unchanged: RAND(expr) and {n,m} path modifiers.
  • npm run util:validateTs reports an error in packages/yasqe/src/tooltip.ts. The same error occurs on main.

🤖 Generated with Claude Code

MathiasVDA and others added 3 commits September 29, 2026 22:01
Recent SWI-Prolog versions treat ==> and => as built-in SSU rule
operators, so grammar productions were compiled as rules instead of
stored as facts, producing an empty table. Store productions as
ebnf/2 and bnf/2 facts, emit ESM directly and format with prettier.
Regenerating the SPARQL 1.1 table yields an identical file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Extend the LL(1) grammar with the SPARQL 1.2 Query Language grammar
(W3C Working Draft of 21 September 2026):

- VERSION declaration
- triple terms <<( s p o )>> in patterns, VALUES and expressions
- reified triples << s p o ~ r >>, reifiers and annotation blocks {| |}
- language tags with base direction ("x"@en--ltr)
- LANGDIR, STRLANGDIR, hasLANG, hasLANGDIR, isTRIPLE, TRIPLE, SUBJECT,
  PREDICATE and OBJECT; aggregates are built-in calls; '!' applies to
  a UnaryExpression

The tokenizer enforces the grammar note that reifiers and annotations
may only follow a simple predicate, not a property path. Also fix
PropertyListPath (a blank node property list may be followed by a
path) and state.complete, which was always false.

Add unit tests with valid and invalid SPARQL 1.2 queries and browser
tests for syntax checking and formatting of SPARQL 1.2 queries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The grammar now covers SPARQL 1.1 and 1.2, so rename
sparql11-grammar.pl to sparql-grammar.pl and gen_sparql11.pl to
gen_sparql.pl, and update build.sh, the README and a code comment.
The "sparql11" mode name is public configuration and stays unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@MathiasVDA MathiasVDA left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated review of this PR: one finding, no correctness bugs. Also checked: the grammar table regenerates with no LL(1) conflicts, the 82 grammar unit tests pass, the simple-predicate rule for reifiers and annotations holds for extra cases (nesting, ;/, lists, multi-line input, inverse paths), the new keywords and punctuation are ordered correctly, and the indentation and build-script changes look fine.

🤖 Generated with Claude Code

'!='= '!=',
'!'= '!',
'<<('= '<<\\(',
'<<'= '<<',

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Compatibility note: with the new << token, a < comparison written without a space before a full IRI is now a syntax error. For example, FILTER(?o<<http://e/a>) used to read as < followed by an IRI and was accepted as SPARQL 1.1. It now reads as << and is flagged. Writing ?o < <http://e/a> still works.

This matches SPARQL 1.2, whose grammar notes say the longest match wins when tokenizing, so a 1.2 parser reads it the same way. Suggest keeping the behaviour and mentioning it in the changeset as a known incompatibility.

A '<' comparison directly followed by a full IRI (?o<<http://...>) is
now a syntax error, because SPARQL 1.2 tokenizes '<<' as a single
token. Mention this in the changeset and add tests for both the
rejected form and the spaced form.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

It rewrites the SPARQL grammar/generator and a ~3000-line machine-generated LL(1) parse table and introduces a documented tokenization breaking change, which cannot be fully verified without executing the Prolog toolchain and warrants human review.

Review effort: Balanced
Findings: None

What changed in this PR

This PR extends the Yasqe SPARQL query editor to understand SPARQL 1.2 syntax (Working Draft of 21 September 2026) in addition to SPARQL 1.1, so 1.2 queries are highlighted and syntax-checked rather than flagged as errors. It works by extending the EBNF grammar (sparql-grammar.pl), regenerating the LL(1) parse table that drives the stream tokenizer, and adding new stateful logic to the tokenizer to enforce the "reifier/annotation only after a simple predicate" rule. The public mode name ("sparql11") is intentionally kept for backward compatibility.

Changes:

  • Adds 1.2 grammar constructs: VERSION declaration, triple terms <<( s p o )>>, reified triples << s p o ~ r >>, reifiers/annotation blocks {| ... |}, base-directional language tags (@en--ltr), new built-ins (LANGDIR, STRLANGDIR, hasLANG, isTRIPLE, TRIPLE, etc.), !-over-UnaryExpression, and aggregates as built-in calls.
  • Fixes the Prolog generator for SWI-Prolog ≥8.3 (stores productions as ebnf/2/bnf/2 facts), fixes two latent SPARQL 1.1 bugs (propertyListPath → propertyListPathNotEmpty; state.complete now actually computed), and renames grammar files to be version-neutral.
  • Adds an 82-case tokenizer unit suite plus browser tests, a changeset noting a <<-tokenization breaking change, and doc updates.
File Description
packages/​yasqe/​grammar/​sparql-grammar.pl Adds SPARQL 1.2 EBNF productions and new terminals/keywords/punctuation.
packages/​yasqe/​grammar/​_tokenizer-table.js Regenerated LL(1) parse table reflecting the new grammar.
packages/​yasqe/​grammar/​tokenizer.ts New verbPaths state + LANG_DIR terminal, multi-char bracket tokens, complete-flag fix.
packages/​yasqe/​grammar/​util/​{gen_ll1,rewrite,prune,ll1,output_to_javascript,gen_sparql}.pl Generator fixes for modern SWI-Prolog (ebnf/2/bnf/2 facts) and ESM output.
packages/​yasqe/​grammar/​build.sh /​ README.md Robust regeneration script (prettier + set -e) and updated docs.
packages/​yasqe/​src/​editor/​language.ts Doc comment updated to mention 1.2.
test/​unit/​yasqe-sparql12-grammar-test.ts New tokenizer-level grammar test suite.
test/​run.ts Browser tests for validity + formatter round-trips.
test/​fix-esm-imports.mjs Copies the generated table into the compiled test tree.
.changeset/​sparql-1-2-syntax.md Minor-version changeset documenting the feature and the breaking change.

I reviewed the hand-written tokenizer state machine (verbPaths tracking, pruneVerbPaths, violatesSimplePathCondition, the complete/$-skipping fix), the grammar/generated-table consistency, the generator rewrite, the renamed-file references, and the tests. The trickiest correctness case (a simple predicate following a complex property path still permitting annotations) is explicitly covered by tests, the file renames have no stale references, and the generated table matches both its .d.ts and its consumers. I found no concrete defects worth an inline comment.


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@MathiasVDA
MathiasVDA merged commit f62052d into main Sep 29, 2026
3 checks passed
@jitsedesmet

Copy link
Copy Markdown

@MathiasVDA Nice progress! If you are looking for more tests on edge cases, you can always chack out whether testing the queries curated by the rdf-tests community working group would help. Similarly, Traqula exposes static queries as curated tests in its own @traqula/test-utils.

Spec tests work like this:
https://github.com/comunica/traqula/blob/c6fd17b03bcfe3014b883b5d932f5a7b039d3d5f/engines/algebra-sparql-1-2/package.json#L42-L46

The Traqula tests have positive tests: https://github.com/comunica/traqula/blob/c6fd17b03bcfe3014b883b5d932f5a7b039d3d5f/engines/parser-sparql-1-2/test/statics.test.ts#L84

and negative tests: https://github.com/comunica/traqula/blob/c6fd17b03bcfe3014b883b5d932f5a7b039d3d5f/engines/parser-sparql-1-2/test/statics.test.ts#L99

(Just an FYI since I got a notification that you implemented 1.2 support :)) )

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add support for SPARQL 1.2

3 participants