Skip to content

mapping_to_ddl.py: measure the sort-key decision instead of defaulting quietly - #45

Merged
litkhai merged 1 commit into
mainfrom
ddl-sort-key
Sep 29, 2026
Merged

litkhai merged 1 commit into
mainfrom
ddl-sort-key

Conversation

@litkhai

@litkhai litkhai commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Closes #39.

The DDL carried ORDER BY (<time field>) and said nothing about it. That runs, passes every
check downstream, and is the one output here that can be quietly wrong: the table works,
costs more to query than it needed to, and the sort key is the most expensive thing to change
once data has landed.

Measure the decision; still refuse to make it

One aggregation request gets cardinality and coverage for every keyword, boolean, ip and
integer field — not text, which is analyzed and is not a sort key:

-- sort key --
  ORDER BY (@timestamp)  <- DEFAULT, not a design
  candidates by cardinality (all 300000 documents):
            4 distinct  100.0% present  log.level
            6 distinct  100.0% present  service.name
            7 distinct  100.0% present  http.response.status_code
           45 distinct  100.0% present  service.version
       296744 distinct  100.0% present  client.ip  (too high to lead a sort key)
       299214 distinct  100.0% present  trace.id  (too high to lead a sort key)

The default ORDER BY is now marked NEEDS REVIEW in the same vocabulary as the columns,
with the measured candidates named inline. The half it cannot measure is which of those
your queries filter on
— --order-by supplies that and is validated against the mapping.

  • --probe-shard-size N measures inside a sampler aggregation for a large index, and a
    distinct count that reaches the sample is reported as "at least", never as a number.
  • --no-probe skips it — and then the DDL says candidates were not measured, rather than
    that none were usable. Those are different statements and conflating them is the failure
    this script exists to refuse.
  • No TTL, no codecs, and the DDL says why: a TTL deletes data, so it needs the retention
    answer rather than a guess.

Measuring it refuted the rule of thumb I was about to write down

Same 300,000 rows, loaded twice into 26.6.8.7, OPTIMIZE FINAL on both:

Sort key On disk
(@timestamp) — the default 22.68 MiB
(log.level, service.name, @timestamp) — the textbook-looking prefix 23.37 MiB

The prefix cost 3% more. This repository's seed is uniformly random, which is the worst
case for a prefix: it destroys the time ordering that compresses timestamps and creates no
runs in exchange. Real observability data is correlated — one service emits runs of one level
— so a prefix usually helps there. Both are in the README, because that is the argument for
deciding from the query pattern rather than from a rule, and it is why the tool prints an
example it explicitly tells you not to take as advice.

Verified on

Elasticsearch 8.17.0 → ClickHouse 26.6.8.7.

Checked Result
ranking against the 300,000-document seed as above, client.ip and trace.id flagged too high
--order-by end to end 4 chunks, 300,000 rows, parity passing, system.tables reporting `log.level`, `service.name`, `@timestamp`
--order-by with an unknown column refused, with the available columns listed
--probe-shard-size 2000 both high-cardinality counts flagged as "at least", the four low ones unchanged
--no-probe DDL says candidates were not measured
an index with no date field ORDER BY tuple() with the note — this crashed on the first attempt, because the report built its example around a time column that was None

AGENTS.md's scope entry for schema design is updated: the sort key is now measured and
flagged, and still not decided for you. That distinction is the whole point of the change.

🤖 Generated with Claude Code

…g quietly

The DDL carried ORDER BY (<time field>) and said nothing about it. That runs,
passes every check downstream, and is the one output here that can be quietly
wrong: the table works, costs more to query than it needed to, and the sort
key is the most expensive thing to change once data has landed.

So the script now measures what the decision depends on and still refuses to
make it. One aggregation request gets cardinality and coverage for every
keyword, boolean, ip and integer field -- not text, which is analyzed and is
not a sort key -- and the report ranks them. The default ORDER BY is marked
NEEDS REVIEW in the same vocabulary as the columns, with the measured
low-cardinality candidates named inline. --order-by supplies the answer and
is validated against the mapping; --probe-shard-size measures inside a
sampler aggregation for a large index, where a distinct count that reaches
the sample is reported as "at least" rather than as a number; --no-probe
skips it, and then the DDL says candidates were not measured rather than that
none were usable.

No TTL and no codecs are emitted, and the DDL says why: a TTL deletes data,
so it needs the retention answer rather than a guess.

Measuring it refuted the rule of thumb I was about to write down. Same
300,000 rows, loaded twice into 26.6.8.7 with OPTIMIZE FINAL on both:
ORDER BY (@timestamp) is 22.68 MiB and
ORDER BY (log.level, service.name, @timestamp) -- the textbook-looking
prefix -- is 23.37 MiB, 3% *more*. This repository's seed is uniformly
random, the worst case for a prefix: it destroys the time ordering that
compresses timestamps and creates no runs in exchange. Real observability
data is correlated and a prefix usually helps there. Both are in the README,
because that is the argument for deciding from the query pattern rather than
from a rule -- and it is why the tool prints an example it tells you not to
take as advice.

Verified on Elasticsearch 8.17.0 into ClickHouse 26.6.8.7: the ranking
against the 300,000-document seed, --order-by applied end to end (300,000
rows, parity passing, system.tables reporting the chosen sorting key),
--order-by rejecting an unknown column with the available ones listed,
--no-probe, --probe-shard-size 2000 flagging both capped counts, and an index
with no date field at all -- which crashed on the first attempt, because the
report built its example around a time column that was None.

Closes #39

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@litkhai
litkhai merged commit a10d676 into main Sep 29, 2026
6 checks passed
@litkhai
litkhai deleted the ddl-sort-key branch September 29, 2026 12:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

mapping_to_ddl.py: the sort key is a default, not a design, and nothing downstream flags that

1 participant