Skip to content

fix(scenarios, simulator): red-team floor, rare quiet lines, unique names, coherence, caller STT and confirmations - #140

Merged
hadarishav merged 10 commits into
devfrom
fix/scenario-names-redteam-coherence
Oct 8, 2026
Merged

hadarishav merged 10 commits into
devfrom
fix/scenario-names-redteam-coherence

Conversation

@KarthikAvinashFI

@KarthikAvinashFI KarthikAvinashFI commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

What

Scenario and simulated-caller quality fixes, found by reading generated suites end to end and by reviewing 60 calls from a production run (transcripts, recordings and evals). All changes are generic; nothing is specific to one agent. Every behaviour change in the diff is listed below with what happens for which input.

Behaviour changes

1. Simulator speech-to-text follows the agent's declared languages

The simulator transcribes the agent (the caller's own speech is generated, not transcribed). It picks one Deepgram nova-3 language per call:

  1. SIMULATOR_STT_LANGUAGE, when set, always wins.
  2. A persona with several languages, or a caller whose language multi covers (en, es, fr, de, hi, ru, pt, ja, it, nl), gets multi. Unchanged from dev.
  3. New: any other caller language gets multi when the agent declares languages that do not include it; otherwise it keeps its own code, as before.
Persona languages Agent languages (from the contract) Deepgram language Change from dev
English (any variant) anything multi none
Spanish, French, Hindi, other covered anything multi none
Arabic not declared ar none
Arabic ["Arabic"] or ["Arabic", "English"] ar none
Arabic ["English"] multi new
Polish, Greek, other uncovered ["English"] multi new
Korean ["Korean"] ko none
Arabic + English anything multi none
none anything multi (via en) none
anything anything, with SIMULATOR_STT_LANGUAGE=ar ar none

Why: an Arabic, Polish or Greek caller talking to an English-only agent had its transcriber set to the caller's language, so the agent's English came back empty and the caller stalled. With multi the caller sees the agent's words and its existing rule ("you understand only the languages you speak") makes it react as someone who does not speak English.

How the agent's languages are read, and what happens to bad values. A new contract field, agent_languages, records the languages the agent's own instructions say it speaks (["English"] for "respond in English only"), and is left empty when they do not say. It travels from contract.json to the call as ALK_AGENT_LANGUAGES. The list is used only when every entry is a known language name or code; otherwise the whole list is ignored and rule 3 does not apply.

agent_languages value Read as Arabic caller gets
English, english, English, en English multi
English, French English, French multi
Arabic, English includes the caller's language ar
Englsh, English (US), English;French, English, Klingon unknown entry, list ignored ar
empty, missing, not a list, unreadable contract.json nothing declared ar

The setting only ever chooses between the caller's own code and multi; it never produces a new code. multi is a Deepgram value, and both production regions configure Deepgram for the simulator's transcriber.

Compatibility. Contracts written before this change have no agent_languages; they load with an empty list and every call behaves exactly as on dev. A single string is accepted as a one-item list. A stale duplicate of the covered-language list (which also listed Arabic, never read at runtime) is removed.

Hearing both English and an uncovered language in one call (an Arabic caller to a bilingual agent) still needs a different STT approach; out of scope here.

2. Simulated caller instructions

Situation in the call Before After
Agent reads a recap and asks for confirmation Caller often read the whole address or summary back Caller confirms the way people do: a yes, or the one detail that is wrong
Caller turn after an agent question Caller sometimes spoke the agent's lines (a recap, "shall I go ahead?") Caller is only ever the caller and never says the agent's lines
Agent's words arrive broken or cut off No rule for it Caller says it did not catch that, in its own language, and lets the agent repeat. No timer, no fixed phrase, so an Arabic caller never says an English "hello"
Agent uses a word "Ask what a word means" applied to ordinary words Applies only to terms a person in the caller's place would not know
Agent answers the caller's question with a general remark, or leaves part of it unanswered, and moves on Caller asked for the missing part once, but only just before closing, so the follow-up was usually lost Caller asks for it once in its next turn, even if the agent has moved on: it answers what the agent asked and asks its own question in the same turn

3. Scenario submit checks (enforced in code, the writer is refused with the reason)

Check Applies when Before After
Red-team floor Slice of 4 or more, the plan deals attack levels, suite is not a growing chat suite No floor; a slice could be all plain scenarios A plain scenario is refused once only the attack slots remain, so at least a fifth (rounded up) of the slice carries an attack
Quiet-line share Suite of 12 or more with a quiet interface level (quiet_line, quiet, clear_line) Interface axis exempt from share limits; one run dealt quiet lines to 40% Quiet level held to a sixth of the slice; other interface levels still exempt
Other level shares Suite of 12 or more One level held to a third Unchanged
Duplicate caller Any submit First name had to be unique across the suite Full name must be unique (case and spacing ignored); a shared first name is fine
Family name Any submit No limit Refused when two other callers in the suite already have that family name

4. Writer and planner guidance (prompt and skill text)

Area After
Coherence New checklist item for writers: who calls, who travels, where they are, the sound around them and what they hold all agree; the caller knows only what someone in their place would; nothing presumes a record the agent's world does not hold
Addresses Every address, a home included, is a real place with its city, in the caller's location
Caller names Common full name, family name the suite does not lean on, nothing famous, fictional or close to a famous name
Scenario names A name mentions a language, accent or condition only when its caller and situation have it. Names carry no counter (-03, -381) or invented code, PIN or reference; a number appears only when it is part of what is tested (terminal-4). Stated in the skill and in the name field of submit_scenario
Use cases real_use_cases are only capabilities the agent's own text states, never stretched from a passing word or a tool name
Planner pre-brief check Red-teaming in every brief, at least a fifth of each writer's scenarios, since a short slice is refused at submit
Run-specific authoring policy Moved from the end of the planner and writer prompts to right after the discovered skills, ahead of the checklists
Guest-booking policy Red-teaming added to the distribution it must not disturb

Sub-goal grading is unchanged.

Verification

Scenario generation, all scenarios read:

Run Agent Size Attacks Quiet lines Notes
before ride booking (prod) 500 7% - first-name rule, TV-character personas, city-less addresses
r28 ride booking 100 17% 7% quiet cap
r29 ride booking 200 18% 6% famous-name echoes remained
r30 ride booking 200 22% 4% invented use case; surnames repeated
r31 ride booking 500 19% 4% real use cases only, no repeated full names, max 2 per family name
r35 ride booking 100 27% 15% final branch
r35 auto insurance 100 23% 16% final branch on an unrelated agent: 7 real use cases, coherent, agent-specific sub-goals
r40 ride booking 200 25% 7% final head: no counters or invented codes in names, no name/persona language mismatch, agent_languages: ["English"] recorded
r41 ride booking 500 11% 15% names clean; attack share lower at 500 because fewer writers were dealt attack levels
r42 admin support 200 15% 16% the agent whose earlier suite named every scenario <prefix>-NN-...: no counters now

Live calls on the r41 environment (3 calls, all booked):

Call Checks Result
Arabic caller, English-only agent agent-language rule (multi) every English agent turn heard; booking completed. Two turns the agent spoke in Arabic came back garbled, since multi does not cover Arabic
Spanish caller covered language, unchanged heard in Spanish and English; booking completed
English, destination corrected at the recap caller confirmation rules caller fixed the one wrong detail, never read the address back or spoke the agent's lines

Live calls for the unanswered-question rule, on a guest ride booking scenario where the caller asks whether its PIN is single-use and the agent replies with a general remark about where the PIN comes from:

Run What the caller did after the general answer Result
Before (rule as on dev) Gave the PIN and moved on to the trip; never asked about PIN validity again Question dropped
After, call 1 Next turn (1:22): asked again whether the PIN is single-use or good for multiple rides, then gave the PIN Asked again in the next turn
After, call 2 Next turn (1:39): asked again; the agent answered that the organization decides, and the caller accepted Asked again in the next turn

Both after-calls went on to a confirmed booking.

Speech-to-text, real call recordings sent to Deepgram nova-3: a Spanish call recovers 270 words with multi (67 with en-US); a call held fully in Arabic recovers 804 words with ar and no Arabic words with multi, which is why uncovered languages keep their own code unless the agent cannot speak them.

Tests

Test Pins
test_the_simulator_transcribes_the_agent_in_the_language_it_will_hear (28 cases) Every row of the two tables in section 1: English variants, covered and uncovered languages, agent lists that include or exclude the caller, case and spacing, codes, misspellings, unknown entries, wrong separators, blank lists, multi-language and empty personas, override
test_the_transcriber_sends_multi_only_where_the_model_covers_it The provider, model and language actually sent for each code
test_the_simulator_definition_reads_the_agents_languages_from_its_settings The setting is read by the simulator definition, ignored when unknown, and loses to the override
test_the_agents_declared_languages_reach_the_simulator_of_every_call End to end through the call runner: the setting survives the environment allowlist and an unrelated variable does not. Verified to fail when the allowlist entry is removed
test_the_agents_declared_languages_reach_the_call_environment Reading contract.json: blanks dropped, missing key, non-list value, malformed JSON, missing file
test_the_contract_records_the_agents_languages_and_older_contracts_still_load Old contracts load with an empty list; a string becomes a one-item list
test_a_slice_that_would_fall_below_the_red_team_floor_refuses_another_plain_scenario, test_the_red_team_floor_leaves_a_tiny_slice_alone Floor refuses at the boundary, accepts attacks, accepts plain while room remains, off without attack levels, off below 4, off without coverage
test_a_quiet_line_is_held_to_a_sixth_of_a_slice Quiet refused past a sixth, allowed below it, other interface levels unaffected
test_callers_are_told_apart_by_full_name_not_first_name, test_a_family_name_is_not_reached_for_a_third_time Shared first name allowed, same full name refused regardless of case and spacing, including names saved by sibling writers; third use of a family name refused
test_the_run_specific_policy_reaches_writers_and_planner_before_their_checklists Policy position in both prompts

Prompt wording (section 2 and 4) is not asserted by string matching. tests/harness failures and errors are identical to dev (environment-dependent tests needing services not available locally).

Linear: TH-8389, TH-8391, TH-8365

…name uniqueness, coherence and real-use-case wording
@KarthikAvinashFI KarthikAvinashFI self-assigned this Oct 1, 2026
@KarthikAvinashFI KarthikAvinashFI changed the title fix(scenarios): red-team floor, rare quiet lines, unique full names, coherence fix(scenarios, simulator): red-team floor, rare quiet lines, unique names, coherence, caller STT and confirmations Oct 2, 2026
@KarthikAvinashFI
KarthikAvinashFI marked this pull request as ready for review October 2, 2026 00:25

@hadarishav hadarishav left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for two reproduced scenario-generation issues, detailed inline: the quiet-line cap rejects edits to existing scenarios, and the new surname limit can be exceeded across parallel writers.

Validation: all 178 tests passed across tests/harness/test_judge.py, tests/harness/test_scenario_source.py, and tests/harness/test_voicemail_call_runner.py. Additional focused reproductions exposed the two gaps. Also reviewed the linked internal-docs #54 language limitations; live calls and generation runs were not repeated.

continue
if not mine or len(levels) < (2 if quiet else 3):
continue
share = max(1, wanted // 6) if quiet else max(1, (wanted + 2) // 3)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Exclude the scenario being replaced from the quiet-line count

The new quiet-line cap also applies when submit_scenario updates an existing scenario, but kept still includes that scenario. With wanted=18 and three existing quiet_line scenarios, resubmitting one of those same names with only a corrected instruction is rejected as already at its whole share, even though the update would leave the quiet count at three. Reproduced through the actual submit_scenario handler: it returns the share refusal before reaching accept_scenario. Exclude the existing scenario by name when calculating the candidate's share, and cover an update at the cap.

for one in kept
if one.persona is not None and one.name != name
}
if sum(1 for one in named if one.split()[-1:] == [family]) >= 2:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Enforce the surname limit across parallel writers

This count sees only this writer's kept plus _first_names_on_disk, which reads saved scenario folders. Parallel _write_slice workers use can_save=False/start_from=[] and append successful submissions to written.jsonl instead, so their callers are invisible to this gate until the suite is saved. Reproduced by journalling Marcus Vance and Philip Vance as sibling submissions: _first_names_on_disk returns an empty set, Heather Vance receives no refusal, and merged retains all three. Read shared submissions for the submit check and enforce the new limit at merge/finalization as well, so concurrent writers cannot collectively exceed two callers per surname.

@KarthikAvinashFI

Copy link
Copy Markdown
Member Author

@hadarishav thanks for the review, both addressed in ea59d08.

  • Quiet-line cap: the share count now skips the scenario being resubmitted, so correcting one at the cap is accepted. Test added for that case.
  • Surname limit: the submit check now also reads written.jsonl, so names a sibling writer journalled count as taken. Test added.

On merge-time enforcement: the parallel path you reproduced through (write_in_parallel / _write_slice in scenarios.py) has no caller in src; the only reference is a test that monkeypatches it. It also carries undefined names (journalled, load_catalogue, write_scenarios, tool, schema, ...), so it would fail if anything called it. In the live path every sub-agent submits through one tool server and one shared kept, with no await between the check and the append, so siblings cannot exceed the limit together. I have not added a merge-time check there; I will remove that dead path in a separate PR instead.

@hadarishav
hadarishav merged commit a0e2ef2 into dev Oct 8, 2026
4 checks passed
@hadarishav
hadarishav deleted the fix/scenario-names-redteam-coherence branch October 8, 2026 10:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants