Skip to content
View yha9806's full-sized avatar
🏠
Working from home
🏠
Working from home

Block or report yha9806

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
yha9806/README.md

Haorui Yu

I am a PhD candidate at the University of Dundee (DJCAD), building reliable agent runtimes and evaluation infrastructure for multimodal and culturally situated AI.

My work connects agent-harness engineering, long-horizon evaluation, evidence governance, and human-review workflows. I care about systems whose runtime behaviour and claims can be inspected—not just demos that appear to work.

vulcaart.art · LinkedIn · Google Scholar · VULCA SDK · DeepSeek Harness infrastructure

Current focus

  • Agent Harness reliability — subagent admission control, lifecycle-safe capacity accounting, cancellation, recovery, and native observability.
  • Multi-agent and long-horizon evaluation — reproducible workloads that expose coordination and resource-control failures.
  • Evidence-governed AI workflows — keeping model output, source support, human review, and release claims distinct.
  • Multimodal cultural evaluation — vision-language models for art, heritage, and visual culture.

Selected work

  • dsh-subagent-admission — an experimental shared lifecycle-admission protocol and reference policy kernel for DeepSeek Harness subagents, with exact-target conformance, crash recovery, and read-only native observability.
  • academic-writing-toolkit — local-first, evidence-controlled workflows for AI-assisted academic writing, bounded revision, review, and release governance.
  • gemini-agent — a local-first Gemini coprocessor for Codex with context packs, diff review, visual gates, and opt-in telemetry.
  • VULCA Visual Control SDK — an agent-native visual editing and evaluation SDK with MCP and ComfyUI integration.
  • VULCA-Bench — a multicultural vision-language benchmark with 7,410 image–critique pairs across eight traditions.

Engineering

I work primarily with Python and TypeScript across MCP and plugin systems, multimodal evaluation pipelines, local-first agent tooling, GitHub Actions, PyPI packaging, and reproducible research infrastructure.

Contact

I am open to Agent Harness, agent infrastructure, and research engineering opportunities, as well as open-source collaboration. The best public way to reach me is LinkedIn.

I also take on scoped paid work in the same areas — a pre-release audit of AI-generated output against a rubric your team signs off on, a custom MCP server or Claude Code plugin, or a bounded creative job taken from brief to a checked, human-released package. Scope and prices are on Contra.

Pinned Loading

  1. dsh-subagent-admission dsh-subagent-admission Public

    Shared lifecycle admission protocol and reference policy kernel for DeepSeek Harness subagents.

    TypeScript 3

  2. academic-writing-toolkit academic-writing-toolkit Public

    Local-first, evidence-controlled academic writing workflows for AI agents, with bounded revision, clean-room review, and release governance.

    Python 32 6

  3. gemini-agent gemini-agent Public

    Local-first Gemini coprocessor for Codex: context packs, diff reviews, visual gates, design drafts, and opt-in telemetry.

    JavaScript 1

  4. VULCA-Bench VULCA-Bench Public

    VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding (7,410 image-critique pairs across 8 traditions)

    Python 1

  5. vulca-org/vulca-visual-control-sdk vulca-org/vulca-visual-control-sdk Public

    Agent-native image-editing SDK for Claude Code. 21 MCP tools + /decompose skill — semantic layer splits, L1–L5 cultural scoring, region inpaint. Powered by ComfyUI, Gemini, or mock.

    Python 15 3