Skip to content
View hk-775's full-sized avatar

Block or report hk-775

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
hk-775/README.md

Harleen Kaur

I build open-source tools for enterprise AI: select the model, govern the action, test the boundary, measure the change.

For AI engineering opportunities and collaboration, contact me on LinkedIn.

My projects address complementary engineering questions. AxonLLM selects models and provider routes. Ostiari governs agent actions. Escape Lab tests containment. Practical Eval Lab evaluates application behavior and regressions.

Together, they support a development loop: route → govern → test containment → evaluate behavior → refine. Evaluation findings guide the next routing, policy, or application change.

Start here → one workflow, its evidence, and the code

Engineering blog · Latest: Designing an Evidence Trail for Agent Actions · RSS

Agent guide · Download public context · Read the repository with GitIngest

Your next step What you will find
Watch A 60-second replay of a synthetic customer-summary workflow, plus a 2:13 AxonLLM operator tour.
Understand The problem, architecture, measured outcome, and limitations.
Inspect Exact source revisions, locked dependencies, runnable commands, and recorded evidence.
Evaluate decisions Documented engineering scope and tradeoffs, with an explicit boundary around leadership and production claims.
Read the blog Engineering questions, implementation choices, measured results, and reproducible evidence.

One body of work

Project The question it answers Explore
AxonLLM Which model and provider should handle this request? Repository · Routing evaluation
Ostiari May this agent take this action, under this policy? Repository · Architecture
Escape Lab Can a prohibited outcome occur despite the controls? Repository · Results and limitations
Practical Eval Lab Did the application improve, and what regressed? Recorded evals and guides · Repository and quickstart

The featured workflow connects AxonLLM, Ostiari, and Escape Lab. Practical Eval Lab is a separate local toolkit for choosing success criteria, comparing candidates, inspecting failures, and applying regression checks.

Evaluate the next change

Practical Eval Lab covers classification, structured extraction, tool calling, RAG, response quality, and multi-step agents. Its tuning webpage, Python/HTTP application adapters, and saved JSON/HTML reports make individual outputs and grading decisions inspectable. Quality gates can catch a regression even when an aggregate score improves.

Explore the evaluation workflow. The hosted viewer opens recorded results; tuning and execution run locally.

Read the evidence in context

The featured example uses real routing and control code with a deterministic provider fixture and synthetic customer data. It is a reproducible integration check. It does not measure production adoption, business savings, live-model quality, or general containment safety.

AxonLLM's separate routing evaluation uses a generated, reviewed test corpus. Its benchmark scores are evaluation results; the operator tour's seeded dashboard values are demonstration data. Neither is presented as customer production evidence.

Practical Eval Lab's bundled candidates use local rules. Its teaching datasets combine synthetic cases with an attributed human-preference sample. The recorded offline results demonstrate evaluation methods; they do not establish live-model quality, production adoption, or safety certification.

Pinned Loading

  1. axonllm axonllm Public

    The neural control plane for enterprise LLMs

    Python

  2. axonllm-ledger axonllm-ledger Public

    Open-source cost intelligence for Amazon Bedrock and AWS AI workloads

    Python

  3. ostiari ostiari Public

    Ostiari — Agent Gateway + Control Plane. Runtime safety layer for AI agents.

    Python 1

  4. OstiariEscapeLab OstiariEscapeLab Public

    Open-source agent-containment benchmark and CI release gate with gVisor isolation, reproducible evidence, incident replay, and AxonLLM integration.

    Python

  5. personalization-control-plane personalization-control-plane Public

    Governed experimentation and recommendation optimization with consent, fairness, approvals, and rollback.

    Python