feat: Implement workflow generation command with YAML validation and LLM integration - #37
Conversation
|
Hey @shreyas-lyzr, pinging you on this one since you opened #9. Quick heads-up on scope: per your follow-up comment on the issue, this PR ships
That said, if you'd prefer to land it all in one go, I'm happy to extend this PR with:
Just let me know which way you want it:
Either way works for me, flagging early so we don't end up doing it twice. |
ee68b20 to
34ff957
Compare
shreyas-lyzr
left a comment
There was a problem hiding this comment.
Three blocking issues and a handful of medium/low concerns. The overall structure is solid — schema-on-disk, injectable LLM, retry loop, good test coverage — but the items below need to be addressed before merge.
shreyas-lyzr
left a comment
There was a problem hiding this comment.
Additional findings in loadFlowDefinition (src/workflows.ts) and its call site in src/voice/server.ts. This function is unchanged in the PR diff but is consumed by the new executeFlow path this PR introduces, so the bugs are activated by this PR.
1. Path traversal — blocking
server.ts:827 constructs the path from user-controlled input:
const flowPath = join(resolve(opts.agentDir), "workflows", flowName + ".yaml");
const flow = await loadFlowDefinition(flowPath);flowName comes from msg.text.match(/@([a-z0-9]+(?:-[a-z0-9]+)*)/) at server.ts:3136, which does constrain characters. But loadFlowDefinition accepts an arbitrary filePath: string and passes it straight to readFile with no bounds check. Any future caller that skips the regex — or any bypass in how chat input reaches this code — leads to arbitrary file read. The function should own its own safety: change its signature to (agentDir: string, flowName: string), validate flowName against the already-present KEBAB_RE, and construct the path internally. The call site becomes loadFlowDefinition(opts.agentDir, flowName).
2. yaml.load() returns non-object values — no explicit type guard
yaml.load() returns null for an empty file and can return a string or number on single-value YAML. The current guard:
if (!data?.name || !data?.steps || !Array.isArray(data.steps))uses optional chaining, so when data is null the condition correctly triggers — but the error says "missing name or steps" rather than "file is empty or not a mapping". More importantly, when data is a string (valid YAML), data?.name is undefined and the same misleading message fires. Add an explicit structural check first:
if (data === null || typeof data !== "object" || Array.isArray(data)) {
throw new Error("Invalid flow definition: file must be a YAML mapping");
}3. description not coerced to string
description: data.description || "" passes through a raw number or boolean if the YAML has description: 42. Any downstream code calling .length or other string methods on it throws. Change to String(data.description || "").
4. Per-step skill not validated for emptiness
String(s.skill || "") maps a missing or null skill to "". executeFlow then builds the prompt:
Use the skill "" (load it with /skill:).
and fires an LLM call with no error. After the steps.map(...), add:
const badStep = steps.findIndex((s) => !s.skill.trim());
if (badStep !== -1) throw new Error(`step[${badStep}] has an empty skill`);5. id / depends_on silently dropped — semantic mismatch with schema
The new workflow.schema.json and WorkflowStep in schemas.ts declare id and depends_on as supported fields. validateWorkflow accepts them. But loadFlowDefinition maps steps to SkillFlowStep which has neither field, so they are silently dropped. executeFlow then runs all steps in declaration order regardless of declared dependencies. A workflow author who writes depends_on based on the schema will get silent wrong behavior. Either: (a) add id?: string; depends_on?: string[] to SkillFlowStep and have executeFlow respect them, or (b) remove those fields from the schema so it doesn't advertise support for them.
…LLM integration - Added `workflow.schema.json` for defining the structure of SkillFlow workflows. - Created `workflow.ts` to handle the `gitclaw workflow generate` command, including parsing flags and invoking the LLM. - Introduced `schemas.ts` for loading and validating workflows against the defined schema. - Developed `workflow-generator.ts` to manage LLM interactions and generate workflows based on user prompts. - Implemented tests for workflow generation and validation to ensure functionality and correctness. - Enhanced `package.json` test script for improved testing capabilities.
…e overwrite option
34ff957 to
6f8f1a4
Compare
|
Thanks for the detailed review, @shreyas-lyzr. All points have been addressed. Summary below. Blocking1.
|
|
|
@krishvsoni Question about skill handling in the generator — does workflow generate only produce the workflow YAML based on the skills it finds installed, or does it also generate the SKILL.md/scripts for any new skill it references? Asking because I tested it locally against a project with 18 real installed skills (none related to weather/SMS), and asked for: "check the weather every morning and text me a summary." It generated this without error: name: morning-weather-summary
Is this expected — i.e. is the user meant to notice the mismatch and create those skills themselves afterward — or should the generator (or the validator) catch this and either refuse/retry when the prompt doesn't match any real installed skill? |
Implemented workflow generation command with YAML validation and LLM integration
Resolves #9
What changed
New files
spec/schemas/workflow.schema.jsonsrc/utils/schemas.tsloadWorkflowSchema(),getWorkflowSchemaText(),validateWorkflow(yamlText) → { valid, errors[], data? }. Hand-rolled validator (no new deps).src/utils/workflow-generator.tsgenerateWorkflow({ prompt, skills, previousWorkflow?, model?, apiKey?, llm? })— builds the system prompt + two few-shot pairs (linear pipeline, approval step), calls the LLM, strips code fences, returns YAML.llmis injectable so tests don't hit the network.src/commands/workflow.tsgitclaw workflow generatewith-d / -p / --refine / -m / --api-key / --dry-run. Runs the 2-retry validation loop and writesworkflows/<slug>.yaml.test/workflow-validator.test.tstest/workflow-generator.test.tstest/ts-resolve-hook.mjsnode --experimental-strip-typescan follow Node16-style.jsinternal imports during tests.Modified files
src/index.tsimport { handleWorkflowCommand }and an early-dispatch branch forgitclaw workflow …(mirrors the existinggitclaw plugin …pattern).package.jsontestscript preloads the resolve hook sonpm testworks out of the box.CLI surface
Why these choices
1. Schema on disk, not inlined
Single source of truth.
spec/schemas/workflow.schema.jsonis loaded at runtime and embedded verbatim into the LLM system prompt — the validator and the generator cannot drift.2. Hand-rolled validator (no AJV)
Zero new deps. AJV isn't in
package.json; rather than addajv+ajv-formats, the validator implements only what this schema needs (type,required,additionalProperties,pattern,minItems,$ref, plus adepends_oncross-field check) and exposes the same{ valid, errors[] }contract.3. LLM via
@mariozechner/pi-ai(injectable)Reuse the existing stack. pi-ai is what gitclaw already uses everywhere — supports all providers, inherits telemetry and cost tracking. The
llm?option makes it swappable, so tests stay offline.4. Retry loop, not best-effort
Bounded self-healing. When the LLM emits invalid YAML, the CLI re-prompts up to 2x with the validator's errors appended; if still invalid it prints errors + the last YAML and
exit(1)— no unbounded loop, no silent failure.5. Argv dispatch, not Commander
Match the codebase. gitclaw has no Commander dep — it uses hand-rolled argv parsing with subcommand dispatch (see
gitclaw plugin …). Theworkflowsubcommand follows the same pattern.6. Flat
requires_approval, not nestedcompliance.*Stay aligned with runtime. gitclaw's existing SkillFlow shape is flat (
skill/prompt/channel). A nestedcomplianceobject would diverge from the loader — the flat field is what the runtime can act on today.Proof
1. Tests pass — 24/24
2. Typecheck clean
3. Pre-existing tests not regressed
Existing
test/telemetry.test.ts— 8/8 still pass under the newnpm testflags.test/sdk.test.tsfails in this environment, but the failure is pre-existing and unrelated: it importsdist/exports.js, which requiresnpm run buildagainst@mariozechner/*packages that aren't installed locally. Nothing in this PR touches the SDK surface.4. End-to-end retry behavior demonstrated in tests
The "retries on invalid YAML, succeeds on second attempt" test plumbs an
LlmClientthat returns invalid YAML once, then valid YAML — and asserts the retry prompt to the LLM included the words "schema validation". From the test output above:The "give up after 2 retries" test:
5. Schema-driven safety: example violations the validator catches
name(root): missing required property "name"name: MyWorkflow(not kebab-case)name: value "MyWorkflow" does not match pattern ^[a-z0-9]+(-[a-z0-9]+)*$steps: []steps: array must have at least 1 item(s), got 0promptsteps[0]: missing required property "prompt"nonsense: truesteps[0]: unknown property "nonsense"depends_on: [does_not_exist]steps[1].depends_on: references unknown step id "does_not_exist"YAML parse error: …Issue → PR mapping (acceptance checklist)
-p "<text>"generateWorkflow+ few-shot examples;depends_onexpresses orderingworkflows/<slug>.yaml--refine <file>modesrc/utils/workflow-generator.tssrc/utils/schemas.ts+ retry loop insrc/commands/workflow.tsNotes for review
tsconfig.json.pi-agent-core+pi-ai) is lazy-imported, so it never loads in tests.OPENAI_API_KEY/<PROVIDER>_API_KEYresolution falls back through--api-key, then provider-specific env, thenOPENAI_API_KEY. Missing key produces a clear error andexit(1).