University of Toronto student intending to pursue a double major in Mathematics and Computer Science, building AI evaluation systems and reliable software products.
I work across Python evaluation infrastructure, TypeScript/React Native products, and backend reliability. I care about reproducible evidence, failure modes, and honest limits—not just demo paths.
Product portfolio · Résumé (PDF) · Email
Expected graduation: May 2030.
Agent Eval Mutation Lab · Python, Inspect, Docker, SQLite
A typed, deterministic engine for testing execution-semantic robustness in tool-agent scorers. It runs 104 canonical tasks with resumable SQLite state, content-addressed evidence, explicit unknown/abstain handling, and clean-checkout artifact reproduction. Its public reports retain failed model-study gates and unfavorable outcomes instead of promoting unsupported conclusions.
- Agent Proof: TypeScript CLI that runs reviewer-selected checks without shell interpolation and writes redacted JSON/Markdown verification reports.
- Microsoft Agent Lightning: merged fixes for shutdown rollout races and local-worker agent URL handling.
- Meridian Labs Inspect Scout: merged per-item model-usage accounting fix with regression coverage.
- Languages: Python, TypeScript, JavaScript, SQL
- AI and evaluation: PyTorch, NumPy, Inspect, structured LLM APIs, deterministic evaluation pipelines
- Web and mobile: React, React Native, Expo, Next.js, Node.js
- Data and reliability: PostgreSQL, Supabase, SQLite, Docker, pytest, mypy, GitHub Actions
- Product integrations: Twilio Voice/SMS, RevenueCat, Vercel
- Turn ambiguous product requirements into tested systems with explicit failure behavior.
- Treat dependencies, logs, provider callbacks, and external inputs as untrusted evidence.
- Separate local or synthetic verification from production, user, revenue, and impact claims.
If you are hiring for a Summer 2027 or selective Fall/Winter AI/software internship, email me.
