Skip to content
View shaurya416's full-sized avatar
🚀
building
🚀
building

Block or report shaurya416

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
shaurya416/README.md

Shaurya Singh · Seattle, WA
cs + data science at UW, class of '27.

everyone sensible says pick a lane, and i keep not doing it. so far that has meant proteins, brains, quantum noise, classrooms, street safety, buildings, an underwater robot, and the code that grades things.

$ grep -rl "same field" ~/built/
no matches

field guide

one line per world i've wandered into.

  • proteins · undergrad research in UW's DAIS lab: preprocessing, training runs and renders for DeepTracer, which turns cryo-EM maps into 3D protein structures. no byline, and that's fine.
  • brains · taught myself neuroscience in public: four essays, optogenetics to connectomics (2022–23). lately, trigalign, for EEG recordings that dropped a trigger or two.
  • quantum · shotfloor: the score a perfect quantum computer gets from shot noise alone. 8 qubits, uniform target, 1,024 shots: 0.92 hellinger fidelity, not 1.
  • classrooms · co-founded SkillTern and was its CTO (2022–23): computer science for 350+ students, built to run without me.
  • street safety · TurtleShell: K-means over 10,000+ LAPD crime datapoints and an iOS SOS app in Swift. got into Microsoft for Startups, then i shut it down on purpose (oct 2023).
  • buildings · a cabin-design prototype that has to pass the Seattle residential code: rule engine, NSGA-II over cost, space and comfort, IFC export. runs locally, so no link.
  • earlier · H2OAquatics, an autonomous underwater vehicle aimed at ocean acidification, the first time the build-at-it reflex showed up. also a co-written policy paper on phishing, runner-up at WatGov (Waterloo, 2023).
  • graders · scoring code that hands out credit it shouldn't. it turned into a side quest; notes at the bottom.
  • off the clock · music. made often, released never. links when that changes, not before.

things i made

small python tools, each about one thing that quietly goes wrong. none is public yet, so no links: each name becomes one the day its repository is.

tests and repos

  • tolwatch · a pytest plugin that flags allclose, isclose and approx checks that can't fail, or are far looser than they need to be.
  • ziptrace · names the zip() and map() calls that silently dropped data because their inputs had different lengths.
  • unadded · finds the untracked or gitignored files a passing test run quietly depends on, before CI does.
  • whymodified · git status says a file is modified and nothing in it changed. this says why, and the one command that clears it.

science data

  • shotfloor · the score a flawless quantum computer gets at a given shot count, and how many shots it needs to reach a target.
  • trigalign · pairs a stimulus log with an EEG trigger list when triggers were dropped or duplicated, or the clocks drift.

grading code

  • scoreprobe · feeds a grader the inputs known to slip through (a blank answer, "not Paris" against "Paris", 4200 against 42) and reports which earned credit.
  • normladder · is "1,000" the same answer as "1000"? 8 of 14 answer-matching rules from lm-eval, HELM, SQuAD and friends say yes. it shows which, and why.
  • mcqlint · lints multiple-choice eval sets: keys that name no option, duplicate options, leaked answers, position bias.
  • tabverify · recomputes the averages, deltas and bold in results tables at the precision they were printed, and flags what no rounding explains.

side quest: grading the graders

the habit: i read the code that decides whether an answer is right (eval harnesses, classroom autograders, forecasting and fairness metrics) and file what i find. since 2026-09-23: 238 issues in 229 repositories, 20 fixed by a merged pull request so far, 17 of those by other people and 3 by me.

one example, small enough to read: scoreprobe on a two-line grader, before and after a one-line fix.

scoreprobe on a grader that returns reference.lower() in answer.lower(): 5 of 9 checks reproduce, for example 'not Paris' earns credit against 'Paris'. With reference.lower() == answer.lower() instead: 0 of 9.

three more panels and the one-defect table

Issues filed per day in UTC, 2026-09-23 to 2026-10-08: 57 44 35 40 4 15 0 38 5 0 0 0 0 0 0 0. 238 filed, 20 fixed by a merged PR. State of each of the 238 issues: fixed by someone else's PR 17, by own PR 3, closed without a merged PR 5, open 213. Median 109.5 hours from report to a merged fix by someone else. 238 issues across 229 repositories in 15 areas; largest: eval harnesses & benchmarks, 42.

eight of the fixes share one defect: a missing, empty or undefined result counted as a score.

report, fixed bywhat was wrong
otter-grader#1020
shaurya416 #1021
grading produced no results and the log still said all tests passed
AReaL#1757
MohammadHijjawi97 #1758
a reply with no answer scored 0.9 against a "None" placeholder gold
evalscope#1783
Yunnglin #1785
a blank answer against a blank reference scored full credit
mlxtend#1202
MohammadHijjawi97 #1203
mcnemar returned p = 0.0 by default when there were no discordant pairs
utilsforecast#285
vgvr0 #286
smape scored a missing forecast as a perfect 0
trl#7364
adithya-s-k #7467
OpenReward's default GRPO reward could not tell never scored from scored 0.0
Curator#2443
1fanwang #2444
an empty document scored 1.0 on the repetition filter and was kept
darts#3221
adigulalkari #3223
grid search returned a NaN-scored combination as the best when it came first in the grid
19 merged pull requests closed 20 of the issues
merged 2026, utc pull request by closes hours
09-24 zjunlp/EasyEdit#717 TengJiao33 #716 4.7
09-25 PrairieLearn/PrairieLearn#15854 jonatanschroeder #15852 26.8
09-25 ucbds-infra/otter-grader#1021 shaurya416 #1020 31.1
09-27 areal-project/AReaL#1758 MohammadHijjawi97 #1757 56.7
09-28 datajuicer/data-juicer#1080 shaurya416 #1077 101.5
09-28 modelscope/evalscope#1785 Yunnglin #1774, #1783 115.0 / 16.6
09-29 Marker-Inc-Korea/AutoRAG#1722 vkehfdl1 #1707 86.7
09-29 rasbt/mlxtend#1203 MohammadHijjawi97 #1202 30.8
09-30 truera/trulens#2836 feiiiiii5 #2823 109.5
09-30 comet-ml/opik#8528 desty #8519 157.7
09-30 terrierteam/ir_measures#84 shaurya416 #83 2.5
10-01 Kaggle/kaggle-benchmarks#220 yyl #217 187.5
10-01 Nixtla/statsforecast#1248 vgvr0 #1247 27.9
10-02 Nixtla/utilsforecast#286 vgvr0 #285 105.9
10-03 EvolvingLMMs-Lab/lmms-eval#1553 longzhenren #1546 235.1
10-06 huggingface/trl#7467 adithya-s-k #7364 296.2
10-07 NVIDIA-NeMo/Curator#2444 1fanwang #2443 316.2
10-07 scikit-learn-contrib/MAPIE#994 SashaMIT #993 236.8
10-08 unit8co/darts#3223 adigulalkari #3221 241.4
method
  • counts are total_count from GitHub search at 2026-10-08 20:50 utc: author:shaurya416 is:public is:issue (238: 213 open, 25 closed) and author:shaurya416 is:public is:pr (33: 3 merged, 28 open, 2 closed unmerged; 33 to other people's repositories), excluding the owner's own organisations.
  • fixed means the issue is closed and a pull request in its repository that references it was merged before it closed; an issue closed as not planned never counts. all 25 closed issues were traced: 20 fixed, 5 closed without a merged pull request. the 16 pull requests by others come from 14 people.
  • an open issue whose fix did not close it is not counted, so fixed is a lower bound.
  • hours run from the issue being filed to the pull request merging.
  • areas: each of the 229 repositories is put in one of 15 areas by hand. that, the four kinds of code the side-quest sentence names, the sentence above the defect table and its one line per row are the hand-written inputs to the side quest.
  • the scoreprobe panel is a captured run of scoreprobe on the code it shows. its first public release is pending; with its source, python3 build.py capture-probe --scoreprobe-src PATH reruns the panel.
  • days are utc. a day with nothing is drawn as a mark, not left blank.
  • regenerate: python3 build.py fetch reads the api (it needs GITHUB_TOKEN), then python3 build.py renders. same input, same bytes.

side-quest numbers: github api, as of 2026-10-08 20:50 utc.

Popular repositories Loading

  1. open-instruct open-instruct Public

    Forked from allenai/open-instruct

    AllenAI's post-training codebase

    Python

  2. SkyRL SkyRL Public

    Forked from NovaSky-AI/SkyRL

    SkyRL: A Modular Full-stack RL Library for LLMs

    Python

  3. slime slime Public

    Forked from THUDM/slime

    slime is an LLM post-training framework for RL Scaling.

    Python

  4. LiveBench LiveBench Public

    Forked from LiveBench/LiveBench

    LiveBench: A Challenging, Contamination-Free LLM Benchmark

    Python

  5. SWE-bench SWE-bench Public

    Forked from SWE-bench/SWE-bench

    SWE-bench: Can Language Models Resolve Real-world Github Issues?

    Python

  6. open-r1 open-r1 Public

    Forked from huggingface/open-r1

    Fully open reproduction of DeepSeek-R1

    Python