Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions labs/elastic-migration/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ and this section applies that rule to the documentation itself.

| Not covered | What that means for you |
|---|---|
| **Schema design.** `mapping_to_ddl.py` emits `ORDER BY (<time field>)` and nothing else: no low-cardinality prefix, no TTL, no codecs, no skip indexes. | The table will be correct and may be far more expensive to query than it needed to be, and the sort key is the one decision that is expensive to change after the data lands. Decide it before the first chunk. [#39](https://github.com/litkhai/clickstack-hyperdx-hols/issues/39) |
| **Schema design.** The sort key is now *measured and flagged* -- `mapping_to_ddl.py` ranks the candidate prefix columns by cardinality and coverage and marks the default `NEEDS REVIEW` -- but it is still not **decided**, because which of those your queries filter on is not in the mapping. No TTL and no codecs are emitted at all. | Give the sort key with `--order-by` once you know the query pattern, and decide the TTL from the retention answer. A low-cardinality prefix is not a free win: measured on this repository's seed it cost 3% more disk than the time column alone (see `data/README.md`). |
| **Scale.** Verified against 300,000 documents (2,000,000 rows for `idmap/`). The ceiling this lab exists to get past is ten million. | The mechanisms are built for scale and were measured below it. Expect to find things at your volume that a 300,000-row run cannot show: PIT keep-alive under load, shard-to-slice ratios, insert pressure. |
| **ClickHouse Cloud specifics.** The `s3()` load, `SSD_CACHE(PATH …)` and the memory ceiling were all measured against a local single node. | The three claims you most want to rely on at real scale are the three never executed against the destination. [#41](https://github.com/litkhai/clickstack-hyperdx-hols/issues/41) |
| **An index pattern whose indices have different mappings.** `plan.py` plans across a pattern; `mapping_to_ddl.py` reads one index and `run.py` loads into one table. | A rollover alias with months of backing indices that each grew their own fields is the normal Elastic case. Check the mappings agree before trusting one table. [#40](https://github.com/litkhai/clickstack-hyperdx-hols/issues/40) |
Expand Down Expand Up @@ -230,7 +230,7 @@ Not a feeling. A list:

| 다루지 않는 것 | 그것이 의미하는 바 |
|---|---|
| **스키마 설계.** `mapping_to_ddl.py`는 `ORDER BY (<시간 필드>)`만 넣습니다. 저카디널리티 접두 컬럼도, TTL도, 코덱도, skip index도 없습니다. | 테이블은 맞게 만들어지지만 필요 이상으로 비싼 쿼리가 될 수 있고, 정렬 키는 데이터가 들어간 뒤에 바꾸기 가장 비싼 결정입니다. 첫 청크 전에 정하세요. [#39](https://github.com/litkhai/clickstack-hyperdx-hols/issues/39) |
| **스키마 설계.** 정렬 키는 이제 **측정되고 표시**됩니다 -- `mapping_to_ddl.py`가 후보 접두 컬럼을 카디널리티·존재 비율로 순위 매기고 기본값에 `NEEDS REVIEW`를 붙입니다 -- 하지만 여전히 **결정되지는** 않습니다. 그중 무엇을 여러분의 쿼리가 필터하는지는 매핑에 없기 때문입니다. TTL과 코덱은 아예 생성하지 않습니다. | 쿼리 패턴을 알게 되면 `--order-by`로 정렬 키를 주고, TTL은 보존 기간 답으로 정하세요. 저카디널리티 접두는 공짜 이득이 아닙니다: 이 저장소 시드에서 측정하니 시간 컬럼만 쓸 때보다 디스크를 3% 더 썼습니다(`data/README.md` 참고). |
| **규모.** 문서 300,000건(`idmap/`은 2,000,000행)으로 검증했습니다. 이 랩이 넘어서려는 천장은 1천만입니다. | 메커니즘은 대규모를 위해 만들었지만 측정은 그 아래에서 했습니다. 30만 행으로는 드러나지 않는 것들 -- 부하 상태의 PIT keep-alive, 샤드 대 슬라이스 비율, INSERT 부하 -- 이 여러분 볼륨에서 나올 수 있습니다. |
| **ClickHouse Cloud 고유 부분.** `s3()` 적재, `SSD_CACHE(PATH …)`, 메모리 상한은 모두 로컬 단일 노드에서 측정했습니다. | 대규모에서 가장 의지하고 싶은 세 주장이 정작 목적지에서 실행되지 않은 셋입니다. [#41](https://github.com/litkhai/clickstack-hyperdx-hols/issues/41) |
| **매핑이 서로 다른 인덱스 패턴.** `plan.py`는 패턴 전체를 계획하지만, `mapping_to_ddl.py`는 인덱스 하나를 읽고 `run.py`는 테이블 하나에 적재합니다. | 각자 필드를 키워온 백킹 인덱스가 몇 달치 쌓인 rollover 별칭은 Elastic에서 예외가 아니라 보통입니다. 테이블 하나를 믿기 전에 매핑이 같은지 확인하세요. [#40](https://github.com/litkhai/clickstack-hyperdx-hols/issues/40) |
Expand Down
113 changes: 113 additions & 0 deletions labs/elastic-migration/data/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -287,6 +287,66 @@ unsupported: 1
(One of those ten is `_id`, which is not part of `_mapping` at all -- it is
synthesized so export/load/parity have a stable row key.)

#### The sort key is a decision, and the script makes you make it

The DDL used to carry `ORDER BY (<time field>)` and say nothing about it. That
runs, passes every check downstream, and is the one output here that can be
*quietly wrong* -- the table works and is more expensive to query than it
needed to be, and the sort key is the most expensive thing to change once the
data has landed.

So the script now measures what the decision depends on and refuses to make
it for you:

```
-- sort key --
ORDER BY (@timestamp) <- DEFAULT, not a design
candidates by cardinality (all 300000 documents):
4 distinct 100.0% present log.level
6 distinct 100.0% present service.name
7 distinct 100.0% present http.response.status_code
45 distinct 100.0% present service.version
296744 distinct 100.0% present client.ip (too high to lead a sort key)
299214 distinct 100.0% present trace.id (too high to lead a sort key)
```

One aggregation request gets cardinality and coverage for every keyword,
boolean, `ip` and integer field -- not `text`, which is analyzed and is not a
sort key. `--probe-shard-size N` measures inside a `sampler` aggregation
instead of over the whole index, and a distinct count that reaches the sample
is reported as *at least* rather than as a number. `--no-probe` skips it, and
then the DDL says candidates were not measured rather than that none were
usable.

**The half it cannot measure is which of those your queries filter on**, and
that is the half that decides. Give it with `--order-by`:

```bash
./mapping_to_ddl.py --index logs-demo --table logs_demo \
--order-by 'log.level, service.name, @timestamp' > ddl.sql
```

**A low-cardinality prefix is not a free win, and measuring it said so.** Same
300,000 rows loaded twice into the pinned 26.6.8.7 target, `OPTIMIZE FINAL`
both:

| Sort key | On disk |
|---|---|
| `(@timestamp)` -- the default | **22.68 MiB** |
| `(log.level, service.name, @timestamp)` -- the textbook-looking prefix | **23.37 MiB** |

The prefix cost 3% *more*. This repository's seed is uniformly random, which
is the worst case for a prefix: it destroys the time ordering that makes
timestamps compress and creates no runs in exchange. Real observability data
is correlated -- one service emits runs of one level -- so a prefix usually
helps there. Both of those are reasons to decide it from the query pattern
and the data rather than from a rule of thumb, which is exactly why the
script prints numbers and an example it tells you not to take as advice.

No `TTL` and no per-column codecs are emitted at all. A TTL deletes data, so
it is not something to guess: it needs the retention answer from
[the parent README](../README.md).

The optional `--manifest` output is the contract with `export.py`: it lists,
per field, whether the raw Elasticsearch value must reach ClickHouse as
nested JSON (`JSON`, `Array(...)` and `Tuple(...)` columns) rather than being
Expand Down Expand Up @@ -828,6 +888,59 @@ unsupported: 1
(이 열 개 중 하나는 `_id`입니다 -- `_mapping`에는 아예 없고,
export·load·정합성 검증이 안정적인 행 키를 갖도록 합성한 것입니다.)

#### 정렬 키는 결정이고, 이 스크립트는 그 결정을 하게 만듭니다

전에는 DDL이 `ORDER BY (<시간 필드>)`를 넣고 그에 대해 아무 말도 하지 않았습니다.
그래도 실행되고, 하위 검사도 모두 통과하며, 여기 있는 출력 중 **조용히 틀릴 수 있는**
유일한 것입니다 -- 테이블은 동작하고 필요 이상으로 비싼 쿼리가 되며, 정렬 키는
데이터가 들어간 뒤에 바꾸기 가장 비싼 것입니다.

그래서 이제 결정에 필요한 것을 **측정**하고, 결정 자체는 대신 하지 않습니다.

```
-- sort key --
ORDER BY (@timestamp) <- DEFAULT, not a design
candidates by cardinality (all 300000 documents):
4 distinct 100.0% present log.level
6 distinct 100.0% present service.name
7 distinct 100.0% present http.response.status_code
45 distinct 100.0% present service.version
296744 distinct 100.0% present client.ip (too high to lead a sort key)
299214 distinct 100.0% present trace.id (too high to lead a sort key)
```

집계 요청 한 번으로 keyword·boolean·`ip`·정수 필드의 카디널리티와 존재 비율을
가져옵니다. `text`는 제외합니다 -- 분석되는 필드이고 정렬 키가 아닙니다.
`--probe-shard-size N`은 인덱스 전체가 아니라 `sampler` 집계 안에서 측정하고, 표본
크기에 도달한 distinct 값은 숫자가 아니라 **최소값**으로 보고합니다. `--no-probe`로
건너뛰면 DDL은 "쓸만한 후보가 없었다"가 아니라 "측정하지 않았다"고 말합니다.

**측정할 수 없는 절반은 그중 무엇을 여러분의 쿼리가 필터하는지**이고, 결정하는 쪽은
그 절반입니다. `--order-by`로 알려주세요.

```bash
./mapping_to_ddl.py --index logs-demo --table logs_demo \
--order-by 'log.level, service.name, @timestamp' > ddl.sql
```

**저카디널리티 접두 컬럼은 공짜 이득이 아니고, 측정이 그렇게 말했습니다.** 같은
300,000행을 고정된 26.6.8.7 목적지에 두 번 적재하고 양쪽 `OPTIMIZE FINAL`:

| 정렬 키 | 디스크 |
|---|---|
| `(@timestamp)` -- 기본값 | **22.68 MiB** |
| `(log.level, service.name, @timestamp)` -- 교과서적으로 좋아 보이는 접두 | **23.37 MiB** |

접두를 넣은 쪽이 3% **더 큽니다**. 이 저장소의 시드는 균일 난수이고, 그것이 접두
컬럼에 최악의 조건입니다 -- 타임스탬프를 압축하게 해주는 시간 순서를 깨뜨리면서 그
대가로 얻는 연속 구간이 없습니다. 실제 관측성 데이터는 상관이 있어서(한 서비스가 같은
레벨을 연달아 내보냄) 접두가 대개 도움이 됩니다. 이 둘 모두가 정렬 키를 경험 법칙이
아니라 **쿼리 패턴과 데이터로** 결정해야 하는 이유이고, 그래서 스크립트가 숫자를
출력하면서 예시는 조언으로 받지 말라고 말합니다.

`TTL`과 컬럼별 코덱은 아예 생성하지 않습니다. TTL은 데이터를 지우므로 추측할 것이
아니라 [상위 README](../README.md)의 보존 기간 답이 필요합니다.

선택적 `--manifest` 출력은 `export.py`와의 계약입니다: 각 필드가 원본
Elasticsearch 값을 점(dot)으로 평탄화된 스칼라 키가 아니라 중첩 JSON
(`JSON`, `Array(...)`, `Tuple(...)` 컬럼)으로 ClickHouse에 전달해야 하는지
Expand Down
Loading
Loading