[FEAT]: Canonical versioned trace/result serializer (WS0-05 PR 1) - #185
Merged
Nina Chikanov (nina-msft) merged 20 commits intoSep 23, 2026
Merged
Nina Chikanov (nina-msft) merged 20 commits into
Nina Chikanov (nina-msft) merged 20 commits into
Conversation
Introduce rampart/core/serialization.py as the single, neutral full-fidelity Result <-> dict round-trip (Decision D6 gate). Every record carries a single root version (rampart.trace.v1) and decoding dispatches on it, failing closed on an unknown major. The canonical layer defines the supported value domain only: enums encode to .value and fail closed on unknown values, harm_category is a passthrough string, floats must be finite, and free-form maps must be JSON-safe. Transport hygiene (ANSI stripping, float normalization, repr() fallback, size caps) stays at the xdist boundary and is not duplicated here. Binary payloads fail closed pending the WS7 artifact resolver rather than being coerced to text. Also lands the written migration policy (docs/concepts/trace-schema.md): additive-optional = no bump, structural = major bump, missing = not recorded, readers fail closed on unknown major, with named reserved additive slots so WS8 provenance needs no hard migration. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
Describe only the canonical serializer behavior present on this branch, express later migration and consumer work as policy constraints, generalize future additions outside the reserved collar fields, and remove the ship gate. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Remove planning references, place constant comments before declarations, drop the unused identity origin field, and make unsupported binary payload messages independent of planned work. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Remove identity and case_id until their producer is implemented, keep the existing pytest attribution fields optional, and make the reserved metadata constant private. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Fail closed on malformed collection fields, incomplete population references, and non-string harm categories. Validate result indices before serialization so the encoder cannot emit boolean indices rejected by the decoder. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Replace handwritten nested codecs with a cached TypeAdapter on Result, keep ResultRecord as the versioned envelope, and generate the open schema with a CI drift check. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Rename the record helpers to serialize_record and deserialize_record, accepting ResultRecord and JSON text respectively. Keep dictionary conversion on the existing record methods. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Restore public dataclass annotations, apply canonical policies through adapter-local schema hooks, document structural schema limits and transport preparation, and add generated round-trip coverage. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Restore types.py to its PR-base contents and supply datetime/Path resolution inside the canonical adapter. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Nina Chikanov (nina-msft)
marked this pull request as ready for review
September 11, 2026 20:46
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Behnam (behnam-o)
previously approved these changes
Sep 22, 2026
Behnam (behnam-o)
left a comment
Contributor
There was a problem hiding this comment.
Overall, this looks pretty good to me. Please take a look at the inline comments. The schema-version verification comment is materially more important than the others: the current drift check can accept a regenerated, breaking v1 schema without requiring TRACE_SCHEMA_VERSION to change. The remaining comments are primarily about keeping this foundation scoped and deferring consumer-specific or speculative behavior.
Bind compatibility declarations to contract content and the PR base, require adjacent version bumps and migration notes for declared breaking changes, and generate version-specific schema files. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve supported metadata through canonical encoding and decoding, reject unsupported values regardless of key, and document shared consumer preparation with loss/truncation preservation. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep general optional-extension and compatibility guidance without committing names or shapes before their producers and consumers exist. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Remove writer readback validation and the Python-mode enum serializer. Preserve Unicode validation and Python ISO timestamps, and document the narrower writer nesting support. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Reject negative record ordinals and invalid population bounds/relationships in the canonical adapter, publish scalar schema constraints, and leave live dataclasses unchanged. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Replace fixed-depth rejection assumptions with controlled Pydantic and JSON boundary failures. Keep supported nested round-trips and document runtime-dependent limits without changing the codec. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Protect earlier schemas on every base comparison while allowing compatible updates to the active major. Cover successive PR decisions, final JSON writer recursion errors, and Pydantic JSON validation errors. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Behnam (behnam-o)
approved these changes
Sep 23, 2026
Behnam (behnam-o)
left a comment
Contributor
There was a problem hiding this comment.
The latest revision addresses my review comments correctly. The compatibility gate now requires an explicit schema-version decision, and the scope, serialization path, metadata behavior, and invariants look good. Focused tests and static checks pass.
Nina Chikanov (nina-msft)
deleted the
nina-msft-trace-result-schema-gate
branch
September 23, 2026 23:58
3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
WS0-05 — PR 1: Canonical versioned trace/result serializer
Establishes the neutral schema API and migration policy. Existing reporting and xdist consumers are not rewired in this PR.
What this lands
rampart/core/serialization.py:serialize_record(*, record: ResultRecord) -> stranddeserialize_record(*, data: str) -> ResultRecord. These functions own envelope construction, body conversion, and version dispatch. NeitherResultnorResultRecordexposesto_dict()/from_dict(). Projections obtain structured canonical data withjson.loads(serialize_record(record=...)); structured transports can supply JSON text to the reader without importing private conversion helpers or maintaining another serializer.TypeAdapter[Result]over the existing standard dataclasses. The adapter and body-schema generation machinery live in the serialization module, not onResult; there is no parallel body model or set of per-type encode/decode functions. Encoding validates the live Result and uses normal Pydantic JSON-mode serialization, including native enum encoding, adapter-local Unicode validation, and Python ISO datetime formatting. The writer does not reconstruct its output through the reader. Nesting limits vary by platform/runtime and can differ between reader and writer: a deeply nested external record may decode but fail to re-encode; limit failures surface asSchemaError.GetPydanticSchemahook configures copies of generated core-schema nodes. Public fields retain plain dictionary/datetime annotations; no class-wide Pydantic configuration or trace validators are attached to the live dataclasses. Other Pydantic adapters retain their normal behavior. The canonical adapter supplies its datetime/Path resolution namespace. Bothresult.pyandtypes.pyare unchanged from the PR base.ResultRecordenvelope — a frozen dataclass referencing the live result and optionalpytest_nodeid/ nonnegativeresult_indexattribution, omitted from output when not recorded. Every serialized record carries the single root versionrampart.trace.v1; unsupported versions fail closed through the decoder registry. The nested result body is not a standalone durable record. Stable identity fields and durable artifact descriptors remain deferred until their implementations exist.[0, 1], without changing live dataclass construction. Python ISO timestamp formatting is preserved. Strings and mapping keys, including attribution, must contain Unicode scalar values: Python surrogate code points are rejected without lossy normalization, while valid JSON surrogate-pair escapes for characters such as emoji remain supported. Malformed JSON, unsupported values, recursion-limit failures, and adapter serializationValueErrors surface asSchemaError.schemas/trace.v1.schema.json.ResultRecord.json_schema()remains the schema-description API, composing the adapter-derived body and versioned envelope with trace payload restrictions, the Request invariant, and ordinal/population numeric bounds. Runtime class docstrings are omitted in favor of concise trace-specific descriptions.scripts/generate_trace_schema.py --checkis the CI drift gate; generation selects the schema filename fromTRACE_SCHEMA_VERSION.schemas/trace-compatibility.jsonbinds the decision and rationale to fingerprints of the models, codec policies, and published schemas. PR CI compares against the target base and requires a same-major compatibility declaration or an adjacent major bump with retained historical schemas and a migration-document reference; the initial contract is declared explicitly. This enforces a recorded decision, not proof of semantic compatibility or migration correctness, which remain review obligations.index < sizeremain decoder requirements. These requirements and runtime nesting limits are documented rather than claimed as JSON Schema guarantees.docs/concepts/trace-schema.md: additive optional fields require no bump; structural changes require a new major, an adjacent upcaster, and an explicit migration API/CLI; absence means not recorded, not a semantic negative. General extension guidance and the release-based support window are documented; future field names and shapes are not reserved. Transport normalization must prepare a separate result before encoding; a binary text placeholder is a lossy transport view, not a durable replay artifact.Coverage
Hypothesis-generated record round-trips through the public JSON APIs cover nested JSON values, all nine nested types, enums, optional fields, timestamps, metadata preservation, output stability, and nonmutation. Example-based coverage includes the removed dictionary-method surface, field exhaustiveness, version dispatch, attribution, ordinal/population invariants on both conversion boundaries, additive tolerance, invalid JSON/types/cycles, independent-adapter isolation, schema semantics, and payload rejection before filesystem access. Additional regressions cover surrogate-containing text and mapping keys, valid Unicode/JSON surrogate-pair escapes, supported nested round-trips, and portable recursion/error-boundary handling with exception chaining. A prepared-copy boundary case demonstrates canonical reuse without modifying or relaxing the original result. Compatibility-gate coverage includes stale declarations, semantic-only input changes, version decisions, migration references, historical-schema preservation, and version-aware generation.
hypothesisandjsonschemaare development-only dependencies.Scope
Reporting projections, xdist/failure-attachment adoption, and stable identity remain separate follow-up implementations. Existing transport coercion and payload-display behavior require explicit preparation and compatibility work, not a mechanical serializer replacement. No public transport-preparation API or artifact representation is introduced here.