Skip to content
View AlexiFeng's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report AlexiFeng

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
alexifeng/README.md

Alexi Feng

I study how AI systems fail—and how evaluations can capture those failures more reliably.

My work focuses on AI evaluation, benchmark design, agent experiments, and reproducible failure analysis. I previously worked on Seed model evaluation at ByteDance and now continue this work as an independent researcher and builder.

Research focus

  • LLM and agent evaluation
  • Benchmark design and evaluation health
  • Failure analysis and reproducible experiments
  • Evaluation tooling and research workflows

Selected research

Writing

I publish research notes in English and Chinese.

Open to collaboration on AI evaluation, benchmark research, and agent reliability.

Pinned Loading

  1. pagefold pagefold Public

    拾页 · Pagefold:按项目整理 Chrome 标签页,保留独立浏览路径,支持可选 AI 分析与可恢复整理。

    JavaScript 1

  2. TaskDock TaskDock Public

    Human-first Todo workbench with attachable Codex and OpenCode sessions

    TypeScript