feat(yasqe): support SPARQL 1.2 syntax - #188
Conversation
Recent SWI-Prolog versions treat ==> and => as built-in SSU rule operators, so grammar productions were compiled as rules instead of stored as facts, producing an empty table. Store productions as ebnf/2 and bnf/2 facts, emit ESM directly and format with prettier. Regenerating the SPARQL 1.1 table yields an identical file. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Extend the LL(1) grammar with the SPARQL 1.2 Query Language grammar
(W3C Working Draft of 21 September 2026):
- VERSION declaration
- triple terms <<( s p o )>> in patterns, VALUES and expressions
- reified triples << s p o ~ r >>, reifiers and annotation blocks {| |}
- language tags with base direction ("x"@en--ltr)
- LANGDIR, STRLANGDIR, hasLANG, hasLANGDIR, isTRIPLE, TRIPLE, SUBJECT,
PREDICATE and OBJECT; aggregates are built-in calls; '!' applies to
a UnaryExpression
The tokenizer enforces the grammar note that reifiers and annotations
may only follow a simple predicate, not a property path. Also fix
PropertyListPath (a blank node property list may be followed by a
path) and state.complete, which was always false.
Add unit tests with valid and invalid SPARQL 1.2 queries and browser
tests for syntax checking and formatting of SPARQL 1.2 queries.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The grammar now covers SPARQL 1.1 and 1.2, so rename sparql11-grammar.pl to sparql-grammar.pl and gen_sparql11.pl to gen_sparql.pl, and update build.sh, the README and a code comment. The "sparql11" mode name is public configuration and stays unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MathiasVDA
left a comment
There was a problem hiding this comment.
Automated review of this PR: one finding, no correctness bugs. Also checked: the grammar table regenerates with no LL(1) conflicts, the 82 grammar unit tests pass, the simple-predicate rule for reifiers and annotations holds for extra cases (nesting, ;/, lists, multi-line input, inverse paths), the new keywords and punctuation are ordered correctly, and the indentation and build-script changes look fine.
🤖 Generated with Claude Code
| '!='= '!=', | ||
| '!'= '!', | ||
| '<<('= '<<\\(', | ||
| '<<'= '<<', |
There was a problem hiding this comment.
Compatibility note: with the new << token, a < comparison written without a space before a full IRI is now a syntax error. For example, FILTER(?o<<http://e/a>) used to read as < followed by an IRI and was accepted as SPARQL 1.1. It now reads as << and is flagged. Writing ?o < <http://e/a> still works.
This matches SPARQL 1.2, whose grammar notes say the longest match wins when tokenizing, so a 1.2 parser reads it the same way. Suggest keeping the behaviour and mentioning it in the changeset as a known incompatibility.
A '<' comparison directly followed by a full IRI (?o<<http://...>) is now a syntax error, because SPARQL 1.2 tokenizes '<<' as a single token. Mention this in the changeset and add tests for both the rejected form and the spaced form. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
It rewrites the SPARQL grammar/generator and a ~3000-line machine-generated LL(1) parse table and introduces a documented tokenization breaking change, which cannot be fully verified without executing the Prolog toolchain and warrants human review.
Review effort: Balanced
Findings: None
What changed in this PR
This PR extends the Yasqe SPARQL query editor to understand SPARQL 1.2 syntax (Working Draft of 21 September 2026) in addition to SPARQL 1.1, so 1.2 queries are highlighted and syntax-checked rather than flagged as errors. It works by extending the EBNF grammar (sparql-grammar.pl), regenerating the LL(1) parse table that drives the stream tokenizer, and adding new stateful logic to the tokenizer to enforce the "reifier/annotation only after a simple predicate" rule. The public mode name ("sparql11") is intentionally kept for backward compatibility.
Changes:
- Adds 1.2 grammar constructs:
VERSIONdeclaration, triple terms<<( s p o )>>, reified triples<< s p o ~ r >>, reifiers/annotation blocks{| ... |}, base-directional language tags (@en--ltr), new built-ins (LANGDIR,STRLANGDIR,hasLANG,isTRIPLE,TRIPLE, etc.),!-over-UnaryExpression, and aggregates as built-in calls. - Fixes the Prolog generator for SWI-Prolog ≥8.3 (stores productions as
ebnf/2/bnf/2facts), fixes two latent SPARQL 1.1 bugs (propertyListPath→propertyListPathNotEmpty;state.completenow actually computed), and renames grammar files to be version-neutral. - Adds an 82-case tokenizer unit suite plus browser tests, a changeset noting a
<<-tokenization breaking change, and doc updates.
| File | Description |
|---|---|
packages/yasqe/grammar/sparql-grammar.pl |
Adds SPARQL 1.2 EBNF productions and new terminals/keywords/punctuation. |
packages/yasqe/grammar/_tokenizer-table.js |
Regenerated LL(1) parse table reflecting the new grammar. |
packages/yasqe/grammar/tokenizer.ts |
New verbPaths state + LANG_DIR terminal, multi-char bracket tokens, complete-flag fix. |
packages/yasqe/grammar/util/{gen_ll1,rewrite,prune,ll1,output_to_javascript,gen_sparql}.pl |
Generator fixes for modern SWI-Prolog (ebnf/2/bnf/2 facts) and ESM output. |
packages/yasqe/grammar/build.sh / README.md |
Robust regeneration script (prettier + set -e) and updated docs. |
packages/yasqe/src/editor/language.ts |
Doc comment updated to mention 1.2. |
test/unit/yasqe-sparql12-grammar-test.ts |
New tokenizer-level grammar test suite. |
test/run.ts |
Browser tests for validity + formatter round-trips. |
test/fix-esm-imports.mjs |
Copies the generated table into the compiled test tree. |
.changeset/sparql-1-2-syntax.md |
Minor-version changeset documenting the feature and the breaking change. |
I reviewed the hand-written tokenizer state machine (verbPaths tracking, pruneVerbPaths, violatesSimplePathCondition, the complete/$-skipping fix), the grammar/generated-table consistency, the generator rewrite, the renamed-file references, and the tests. The trickiest correctness case (a simple predicate following a complex property path still permitting annotations) is explicitly covered by tests, the file renames have no stale references, and the generated table matches both its .d.ts and its consumers. I found no concrete defects worth an inline comment.
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
@MathiasVDA Nice progress! If you are looking for more tests on edge cases, you can always chack out whether testing the queries curated by the rdf-tests community working group would help. Similarly, Traqula exposes static queries as curated tests in its own Spec tests work like this: The Traqula tests have positive tests: https://github.com/comunica/traqula/blob/c6fd17b03bcfe3014b883b5d932f5a7b039d3d5f/engines/parser-sparql-1-2/test/statics.test.ts#L84 and negative tests: https://github.com/comunica/traqula/blob/c6fd17b03bcfe3014b883b5d932f5a7b039d3d5f/engines/parser-sparql-1-2/test/statics.test.ts#L99 (Just an FYI since I got a notification that you implemented 1.2 support :)) ) |
Fixes #2
Adds SPARQL 1.2 support to the Yasqe query editor, so SPARQL 1.2 queries are highlighted and syntax checked instead of being flagged as errors. The grammar follows the SPARQL 1.2 Query Language Working Draft of 21 September 2026.
What's supported now
VERSION "1.2"declarations<<( s p o )>>in triple patterns,VALUESand expressions<< s p o ~ r >>, reifiers (~) and annotation blocks{| ... |}"hello"@en--ltrLANGDIR,STRLANGDIR,hasLANG,hasLANGDIR,isTRIPLE,TRIPLE,SUBJECT,PREDICATEandOBJECT!applied to any unary expression (!!BOUND(?x)), and aggregates wherever a built-in call is allowed (e.g.ORDER BY COUNT(?x))The spec's grammar notes only allow a reifier or annotation after a simple predicate (an IRI,
aor a variable). The editor now reports this as an error when a property path such as:p/:qcomes first, with an explanatory tooltip.Commits
==>and=>are built-in rule operators. The generator therefore produced an empty table, so the grammar could not be changed. Productions are now stored asebnf/2/bnf/2facts. Regenerating the existing SPARQL 1.1 grammar with the fixed generator gives a byte-identical_tokenizer-table.js.[ :p ?o ] :a/:b ?cwas reported as a syntax error. It is valid SPARQL 1.1.state.completeflag was alwaysfalse. Nothing in Yasqe read it, but the new tests use it.sparql11-grammar.pl→sparql-grammar.plandgen_sparql11.pl→gen_sparql.pl. The public"sparql11"mode name is unchanged.Testing
test/unit/yasqe-sparql12-grammar-test.tswith 82 cases:test/run.ts:npm run unit-test: 154 passing.npm run puppeteer-test: 40 passing.Known incompatibility
A
<comparison written without a space before a full IRI, such asFILTER(?o<<http://example.org/a>), used to be accepted and is now reported as a syntax error. SPARQL 1.2 reads<<as a single token (the longest match wins), so a SPARQL 1.2 parser rejects this query as well. Adding a space fixes it:FILTER(?o < <http://example.org/a>). This is also noted in the changeset.Notes for reviewers
DELETE DATA,DELETE WHEREandDELETEclauses. That is enforced for explicit forms (_:b,[],~ _:r). A reified triple without an explicit reifier implicitly creates a blank node; it is still allowed there, because the spec text only covers explicit syntax.RAND(expr)and{n,m}path modifiers.npm run util:validateTsreports an error inpackages/yasqe/src/tooltip.ts. The same error occurs onmain.🤖 Generated with Claude Code