weights → agents → learning loops
Founder @ xDAN AI. I used to train LLMs. Now I build agent-native products and systems. Next, I want to close the loop: turn agent experience into better models.
Homepage · Hugging Face · GitHub
Worked on the Cloud OS team. My background starts with the systems underneath the applications.
Led our team's model training work at xDAN AI, spanning training and post-training. Public model releases are linked below.
Building products and systems around agents that can use tools and execute real tasks. Moving from model capabilities to working applications.
Exploring metaRSI and how agent experience could feed back into model training. The question I want to work on: can each cycle produce a better model, with improvements that hold up on independent evaluations?
Public artifacts from our team at xDAN-AI:
| Model | Release |
|---|---|
| xDAN-L1-Chat-RL-v1 | 7B model trained with SFT and DPO. |
| xDAN-L3-MoE-Performance-RLHF-0416 | 141B mixture-of-experts model. |
| xDAN-L2-Qwen2.5-32b-Instruct-RM | 32B model release. |



