Repository navigation
fix(scenarios, simulator): red-team floor, rare quiet lines, unique names, coherence, caller STT and confirmations - #140
Conversation
…name uniqueness, coherence and real-use-case wording
…read-backs, stay in the caller role
…the caller's language (TH-8389)
…s speech is cut off
…wiring and scenario submit checks
… a general answer
hadarishav
left a comment
There was a problem hiding this comment.
Requesting changes for two reproduced scenario-generation issues, detailed inline: the quiet-line cap rejects edits to existing scenarios, and the new surname limit can be exceeded across parallel writers.
Validation: all 178 tests passed across tests/harness/test_judge.py, tests/harness/test_scenario_source.py, and tests/harness/test_voicemail_call_runner.py. Additional focused reproductions exposed the two gaps. Also reviewed the linked internal-docs #54 language limitations; live calls and generation runs were not repeated.
| continue | ||
| if not mine or len(levels) < (2 if quiet else 3): | ||
| continue | ||
| share = max(1, wanted // 6) if quiet else max(1, (wanted + 2) // 3) |
There was a problem hiding this comment.
[P2] Exclude the scenario being replaced from the quiet-line count
The new quiet-line cap also applies when submit_scenario updates an existing scenario, but kept still includes that scenario. With wanted=18 and three existing quiet_line scenarios, resubmitting one of those same names with only a corrected instruction is rejected as already at its whole share, even though the update would leave the quiet count at three. Reproduced through the actual submit_scenario handler: it returns the share refusal before reaching accept_scenario. Exclude the existing scenario by name when calculating the candidate's share, and cover an update at the cap.
| for one in kept | ||
| if one.persona is not None and one.name != name | ||
| } | ||
| if sum(1 for one in named if one.split()[-1:] == [family]) >= 2: |
There was a problem hiding this comment.
[P2] Enforce the surname limit across parallel writers
This count sees only this writer's kept plus _first_names_on_disk, which reads saved scenario folders. Parallel _write_slice workers use can_save=False/start_from=[] and append successful submissions to written.jsonl instead, so their callers are invisible to this gate until the suite is saved. Reproduced by journalling Marcus Vance and Philip Vance as sibling submissions: _first_names_on_disk returns an empty set, Heather Vance receives no refusal, and merged retains all three. Read shared submissions for the submit check and enforce the new limit at merge/finalization as well, so concurrent writers cannot collectively exceed two callers per surname.
… journalled caller names
|
@hadarishav thanks for the review, both addressed in ea59d08.
On merge-time enforcement: the parallel path you reproduced through ( |
What
Scenario and simulated-caller quality fixes, found by reading generated suites end to end and by reviewing 60 calls from a production run (transcripts, recordings and evals). All changes are generic; nothing is specific to one agent. Every behaviour change in the diff is listed below with what happens for which input.
Behaviour changes
1. Simulator speech-to-text follows the agent's declared languages
The simulator transcribes the agent (the caller's own speech is generated, not transcribed). It picks one Deepgram nova-3 language per call:
SIMULATOR_STT_LANGUAGE, when set, always wins.multicovers (en, es, fr, de, hi, ru, pt, ja, it, nl), getsmulti. Unchanged fromdev.multiwhen the agent declares languages that do not include it; otherwise it keeps its own code, as before.devmultimultiar["Arabic"]or["Arabic", "English"]ar["English"]multi["English"]multi["Korean"]komultimulti(viaen)SIMULATOR_STT_LANGUAGE=ararWhy: an Arabic, Polish or Greek caller talking to an English-only agent had its transcriber set to the caller's language, so the agent's English came back empty and the caller stalled. With
multithe caller sees the agent's words and its existing rule ("you understand only the languages you speak") makes it react as someone who does not speak English.How the agent's languages are read, and what happens to bad values. A new contract field,
agent_languages, records the languages the agent's own instructions say it speaks (["English"]for "respond in English only"), and is left empty when they do not say. It travels fromcontract.jsonto the call asALK_AGENT_LANGUAGES. The list is used only when every entry is a known language name or code; otherwise the whole list is ignored and rule 3 does not apply.agent_languagesvalueEnglish,english,English,enmultiEnglish, FrenchmultiArabic, EnglisharEnglsh,English (US),English;French,English, Klingonarcontract.jsonarThe setting only ever chooses between the caller's own code and
multi; it never produces a new code.multiis a Deepgram value, and both production regions configure Deepgram for the simulator's transcriber.Compatibility. Contracts written before this change have no
agent_languages; they load with an empty list and every call behaves exactly as ondev. A single string is accepted as a one-item list. A stale duplicate of the covered-language list (which also listed Arabic, never read at runtime) is removed.Hearing both English and an uncovered language in one call (an Arabic caller to a bilingual agent) still needs a different STT approach; out of scope here.
2. Simulated caller instructions
3. Scenario submit checks (enforced in code, the writer is refused with the reason)
quiet_line,quiet,clear_line)4. Writer and planner guidance (prompt and skill text)
-03,-381) or invented code, PIN or reference; a number appears only when it is part of what is tested (terminal-4). Stated in the skill and in thenamefield ofsubmit_scenarioreal_use_casesare only capabilities the agent's own text states, never stretched from a passing word or a tool nameSub-goal grading is unchanged.
Verification
Scenario generation, all scenarios read:
agent_languages: ["English"]recorded<prefix>-NN-...: no counters nowLive calls on the r41 environment (3 calls, all booked):
multi)multidoes not cover ArabicLive calls for the unanswered-question rule, on a guest ride booking scenario where the caller asks whether its PIN is single-use and the agent replies with a general remark about where the PIN comes from:
dev)Both after-calls went on to a confirmed booking.
Speech-to-text, real call recordings sent to Deepgram nova-3: a Spanish call recovers 270 words with
multi(67 withen-US); a call held fully in Arabic recovers 804 words witharand no Arabic words withmulti, which is why uncovered languages keep their own code unless the agent cannot speak them.Tests
test_the_simulator_transcribes_the_agent_in_the_language_it_will_hear(28 cases)test_the_transcriber_sends_multi_only_where_the_model_covers_ittest_the_simulator_definition_reads_the_agents_languages_from_its_settingstest_the_agents_declared_languages_reach_the_simulator_of_every_calltest_the_agents_declared_languages_reach_the_call_environmentcontract.json: blanks dropped, missing key, non-list value, malformed JSON, missing filetest_the_contract_records_the_agents_languages_and_older_contracts_still_loadtest_a_slice_that_would_fall_below_the_red_team_floor_refuses_another_plain_scenario,test_the_red_team_floor_leaves_a_tiny_slice_alonetest_a_quiet_line_is_held_to_a_sixth_of_a_slicetest_callers_are_told_apart_by_full_name_not_first_name,test_a_family_name_is_not_reached_for_a_third_timetest_the_run_specific_policy_reaches_writers_and_planner_before_their_checklistsPrompt wording (section 2 and 4) is not asserted by string matching.
tests/harnessfailures and errors are identical todev(environment-dependent tests needing services not available locally).Linear: TH-8389, TH-8391, TH-8365