I am a PhD candidate at the University of Dundee (DJCAD), building reliable agent runtimes and evaluation infrastructure for multimodal and culturally situated AI.
My work connects agent-harness engineering, long-horizon evaluation, evidence governance, and human-review workflows. I care about systems whose runtime behaviour and claims can be inspected—not just demos that appear to work.
vulcaart.art · LinkedIn · Google Scholar · VULCA SDK · DeepSeek Harness infrastructure
- Agent Harness reliability — subagent admission control, lifecycle-safe capacity accounting, cancellation, recovery, and native observability.
- Multi-agent and long-horizon evaluation — reproducible workloads that expose coordination and resource-control failures.
- Evidence-governed AI workflows — keeping model output, source support, human review, and release claims distinct.
- Multimodal cultural evaluation — vision-language models for art, heritage, and visual culture.
- dsh-subagent-admission — an experimental shared lifecycle-admission protocol and reference policy kernel for DeepSeek Harness subagents, with exact-target conformance, crash recovery, and read-only native observability.
- academic-writing-toolkit — local-first, evidence-controlled workflows for AI-assisted academic writing, bounded revision, review, and release governance.
- gemini-agent — a local-first Gemini coprocessor for Codex with context packs, diff review, visual gates, and opt-in telemetry.
- VULCA Visual Control SDK — an agent-native visual editing and evaluation SDK with MCP and ComfyUI integration.
- VULCA-Bench — a multicultural vision-language benchmark with 7,410 image–critique pairs across eight traditions.
I work primarily with Python and TypeScript across MCP and plugin systems, multimodal evaluation pipelines, local-first agent tooling, GitHub Actions, PyPI packaging, and reproducible research infrastructure.
I am open to Agent Harness, agent infrastructure, and research engineering opportunities, as well as open-source collaboration. The best public way to reach me is LinkedIn.
I also take on scoped paid work in the same areas — a pre-release audit of AI-generated output against a rubric your team signs off on, a custom MCP server or Claude Code plugin, or a bounded creative job taken from brief to a checked, human-released package. Scope and prices are on Contra.




