Skip to content

elastic-migration: authenticate to Elasticsearch (the lab could not connect to a real 8.x cluster) - #43

Merged
litkhai merged 1 commit into
mainfrom
es-auth
Sep 29, 2026
Merged

litkhai merged 1 commit into
mainfrom
es-auth

Conversation

@litkhai

@litkhai litkhai commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Closes #42. Found while answering "can another agent just follow this lab?" — the honest
answer was no, and this is the first reason why.

Every Elasticsearch request in the lab was unauthenticated. The only two Authorization
headers in data/ were both on the ClickHouse side, and _base/'s throwaway cluster runs
with xpack.security.enabled=false — so the tools worked there and could not connect to any
real 8.x cluster, where security is on by default. Four tools, all failing at their first
request.

es_client.py, shared rather than copied four times

The credential handling is the part that is either right or wrong in exactly one place:

  • basic auth (ES_USER/ES_PASSWORD) or an API key (ES_API_KEY), and both at once
    is an error, not a precedence rule — a tool that silently picks one gets debugged
    against the wrong identity.
  • TLS against a private CA (ES_CA_CERT), because 8.x generates its own on first start.
    Plus --es-insecure, which warns on every call: the alternative to that flag is not
    "everyone configures a CA properly", it is "someone disables security on the source
    cluster".
  • Credentials come from the environment and are refused in a URL — a URL ends up in shell
    history, in ps and in error messages. run.py passes them to the export.py it spawns
    through the child's environment, never its argv, for the same reason.
  • 401 and 403 are translated into the thing to change, and a wrong password prints a
    sentence instead of a traceback.

The minimum privileges, found by narrowing a key until each tool broke

{"cluster": ["monitor"],
 "index": [{"names": ["logs-*"], "privileges": ["read", "view_index_metadata", "monitor"]}]}

The index-level monitor is the one that gets left out. _cat/indices needs both
cluster:monitor/state and indices:monitor/stats, and the 403 names the action rather
than the privilege — so the hint says it outright. My first attempt at that hint was wrong
in exactly this way until the cluster corrected it.

A reproducible home for the authenticated path

_base/ gains an elastic-secure profile: the same Elasticsearch with security on, port
9201. HTTP TLS is off there on purpose — authentication was the gap, and a self-signed CA on
top would mean copying a certificate out of a container before anything runs, which makes a
profile nobody uses. The CA path is verified against a default-configuration container
instead.

Verified on

Elasticsearch 8.17.0 with xpack.security.enabled=true, plus a default-configuration 8.17.0
container over https with its generated CA.

Checked Result
whole pipeline, basic auth seed → plan.py → mapping_to_ddl.py → run.py → parity_checks.py, 20,000 documents, 3/3 chunks verified
whole pipeline, API key scoped to exactly those privileges same, 20,000 rows / 20,000 distinct _id in ClickHouse 26.6.8.7
https with ES_CA_CERT connects; the summary line names the CA it verified against
https without it CERTIFICATE_VERIFY_FAILED → one line of cause, one of remedy
--es-insecure runs, warns every call
API key and basic auth refused
credentials in the URL refused, URL redacted in the message
the unauthenticated elastic profile still end to end: 300,000 documents, 4 chunks, parity passing — the quick path did not pay for this

Two bugs this caught in my own work, both only visible by running it: parity_checks.py got
the new flags and never applied them (401 against a cluster it had credentials for), and the
401 hint disappeared once the error was wrapped in a RuntimeError — the status code was in
__cause__.

Remaining gaps from the same review are tracked separately: #39 (the sort key is a default,
not a design), #40 (a pattern of indices with different mappings), #41 (the Cloud path is
documented but never run).

🤖 Generated with Claude Code

Every Elasticsearch request in the lab was unauthenticated. The only two
Authorization headers in data/ were both on the ClickHouse side, and the
throwaway cluster in _base/ runs with xpack.security.enabled=false -- so the
tools worked there and could not connect to any real 8.x cluster, where
security is on by default. Found while answering "can another agent just
follow this?", which is the honest answer to that question.

es_client.py is shared rather than copied into four scripts, because the
credential handling is the part that is either right or wrong in one place:

  * basic auth (ES_USER/ES_PASSWORD) or an API key (ES_API_KEY), and both at
    once is an error rather than a precedence rule -- a tool that silently
    picks one gets debugged against the wrong identity
  * TLS against a private CA (ES_CA_CERT), since 8.x generates its own on
    first start, plus --es-insecure that warns on every call. The alternative
    to that flag is not "everyone configures a CA properly", it is "someone
    disables security on the source cluster"
  * credentials come from the environment and are refused in a URL, which
    ends up in shell history, ps and error messages. run.py passes them to
    the export.py it spawns in the child's environment, never in its argv
  * 401 and 403 are translated into the thing to change, and a wrong password
    prints a sentence instead of a traceback

The minimum privileges, found by narrowing an API key until each tool broke:
cluster [monitor] and index [read, view_index_metadata, monitor]. The
index-level monitor is the one that gets left out -- _cat/indices needs both
cluster:monitor/state and indices:monitor/stats, and the 403 names the action
rather than the privilege, which is why the hint says it outright.

_base/ gains an `elastic-secure` profile: the same Elasticsearch with security
on, on 9201, so the authenticated path has a reproducible home instead of
being tested once by hand. HTTP TLS is off there on purpose -- authentication
was the gap, and a self-signed CA on top would mean copying a certificate out
of a container before anything runs at all. The CA path is verified against a
default-configuration container instead.

Verified on Elasticsearch 8.17.0 with security enabled: the whole pipeline
with basic auth, then again with an API key scoped to exactly those
privileges (20,000 documents, 3 chunks, 20,000 distinct _id in ClickHouse
26.6.8.7); https with ES_CA_CERT connects, without it fails with one line of
cause and one of remedy, and --es-insecure runs while warning. Both
credentials together and credentials in a URL are refused. The unauthenticated
path still runs end to end -- 300,000 documents, 4 chunks, parity passing --
so the quick path did not pay for this.

Closes #42

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@litkhai
litkhai merged commit 6285d79 into main Sep 29, 2026
6 checks passed
@litkhai
litkhai deleted the es-auth branch September 29, 2026 11:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Every Elasticsearch request is unauthenticated: the lab cannot connect to a real 8.x cluster

1 participant