AI safety research, infrastructure, and BioML. Currently a research intern at EleutherAI (SOAR), working on cheap evaluation methods for inoculation prompting. In terms of engineering I build data pipelines, backend services, and ML systems for AI safety and BioML research.
San Diego, CA · LinkedIn · Email
Inoculation prompt evaluation — EleutherAI SOAR. A low-cost heuristic for predicting when a training-time inoculation prompt suppresses a target trait without expensive finetuning runs. Some skills/tools: Fine-tuning infrastructure, LLM-as-judge evaluation pipeline, Tinker API and cloud GPUs.
UmamiBench — a benchmark measuring scientific judgment in frontier models. Some skills/tools: Construct definition, task rubric design, and the model evaluation pipeline.
HELIX — ML architectures for identifying selective RecA antibiotic adjuvants that reduce quinolone resistance. Designed the full computational stack: a docking pipeline over millions of compounds (20 → 2,500 ligands/hr via sharding, caching, and parallelization; 40k+ CPU hours on GCP), custom potency classifiers, and dataset curation.
Some current safety work isn't public yet.
Selected merged pull requests from @OpenCodingSociety:
| PR | What it does |
|---|---|
spring#78 |
General S3 file API — artifact upload/retrieval service, credential handling, storage abstraction |
spring#108 |
S3 integration for the analytics layer |
flask#38 |
Backend API endpoints |
pages#409 |
LLM integration over platform analytics data, admin stats API, admin-only access control |
pages#406 |
API integration across the frontend |
pages#630 |
Admin analytics dashboard |
pages#1011 |
Frontend feature work |
Mechanistic interpretability and evaluation · Empirical Alignment and Pretraining · ML for biology and scientific imaging · Data pipelines and backend infrastructure



