mapping_to_ddl.py: measure the sort-key decision instead of defaulting quietly - #45
Merged
Merged
Conversation
…g quietly The DDL carried ORDER BY (<time field>) and said nothing about it. That runs, passes every check downstream, and is the one output here that can be quietly wrong: the table works, costs more to query than it needed to, and the sort key is the most expensive thing to change once data has landed. So the script now measures what the decision depends on and still refuses to make it. One aggregation request gets cardinality and coverage for every keyword, boolean, ip and integer field -- not text, which is analyzed and is not a sort key -- and the report ranks them. The default ORDER BY is marked NEEDS REVIEW in the same vocabulary as the columns, with the measured low-cardinality candidates named inline. --order-by supplies the answer and is validated against the mapping; --probe-shard-size measures inside a sampler aggregation for a large index, where a distinct count that reaches the sample is reported as "at least" rather than as a number; --no-probe skips it, and then the DDL says candidates were not measured rather than that none were usable. No TTL and no codecs are emitted, and the DDL says why: a TTL deletes data, so it needs the retention answer rather than a guess. Measuring it refuted the rule of thumb I was about to write down. Same 300,000 rows, loaded twice into 26.6.8.7 with OPTIMIZE FINAL on both: ORDER BY (@timestamp) is 22.68 MiB and ORDER BY (log.level, service.name, @timestamp) -- the textbook-looking prefix -- is 23.37 MiB, 3% *more*. This repository's seed is uniformly random, the worst case for a prefix: it destroys the time ordering that compresses timestamps and creates no runs in exchange. Real observability data is correlated and a prefix usually helps there. Both are in the README, because that is the argument for deciding from the query pattern rather than from a rule -- and it is why the tool prints an example it tells you not to take as advice. Verified on Elasticsearch 8.17.0 into ClickHouse 26.6.8.7: the ranking against the 300,000-document seed, --order-by applied end to end (300,000 rows, parity passing, system.tables reporting the chosen sorting key), --order-by rejecting an unknown column with the available ones listed, --no-probe, --probe-shard-size 2000 flagging both capped counts, and an index with no date field at all -- which crashed on the first attempt, because the report built its example around a time column that was None. Closes #39 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #39.
The DDL carried
ORDER BY (<time field>)and said nothing about it. That runs, passes everycheck downstream, and is the one output here that can be quietly wrong: the table works,
costs more to query than it needed to, and the sort key is the most expensive thing to change
once data has landed.
Measure the decision; still refuse to make it
One aggregation request gets cardinality and coverage for every keyword, boolean,
ipandinteger field — not
text, which is analyzed and is not a sort key:The default
ORDER BYis now markedNEEDS REVIEWin the same vocabulary as the columns,with the measured candidates named inline. The half it cannot measure is which of those
your queries filter on —
--order-bysupplies that and is validated against the mapping.--probe-shard-size Nmeasures inside asampleraggregation for a large index, and adistinct count that reaches the sample is reported as "at least", never as a number.
--no-probeskips it — and then the DDL says candidates were not measured, rather thanthat none were usable. Those are different statements and conflating them is the failure
this script exists to refuse.
TTL, no codecs, and the DDL says why: a TTL deletes data, so it needs the retentionanswer rather than a guess.
Measuring it refuted the rule of thumb I was about to write down
Same 300,000 rows, loaded twice into 26.6.8.7,
OPTIMIZE FINALon both:(@timestamp)— the default(log.level, service.name, @timestamp)— the textbook-looking prefixThe prefix cost 3% more. This repository's seed is uniformly random, which is the worst
case for a prefix: it destroys the time ordering that compresses timestamps and creates no
runs in exchange. Real observability data is correlated — one service emits runs of one level
— so a prefix usually helps there. Both are in the README, because that is the argument for
deciding from the query pattern rather than from a rule, and it is why the tool prints an
example it explicitly tells you not to take as advice.
Verified on
Elasticsearch 8.17.0 → ClickHouse 26.6.8.7.
client.ipandtrace.idflagged too high--order-byend to endsystem.tablesreporting`log.level`, `service.name`, `@timestamp`--order-bywith an unknown column--probe-shard-size 2000--no-probeORDER BY tuple()with the note — this crashed on the first attempt, because the report built its example around a time column that wasNoneAGENTS.md's scope entry for schema design is updated: the sort key is now measured andflagged, and still not decided for you. That distinction is the whole point of the change.
🤖 Generated with Claude Code