AI Safety Researcher Β· Research Engineer Β· UC Berkeley
I'm a UC Berkeley student studying Applied Mathematics, Data Science, and Computer Science. I build evaluation systems for language models and tool-using agents, with a focus on whether safety behavior survives deployment pressure.
My focus: the gap between an AI system that looks safe on average and one whose individual decisions remain dependable under real-world constraints.
- Avocado (private research collaboration) β leading the durability track, testing corrigibility and jailbreak durability in fine-tuned LLMs.
- Safety Invariance β measuring whether FP16, INT8, and NF4 quantization preserve the safety decisions of tool-using agents.
- Anketa β building Python NLP pipelines, REST microservices, and Swift-based iOS features as a software engineering intern.
I presented first-authored research at Stanford's PAI26 Conference on the epistemic safety of AI-generated physics explanations.
The pilot study found that AI explanations matched human explanations on correctness, while showing lower completeness and less intuition-bridging content. View the research, paper, and poster β
| Project | Stack | Result |
|---|---|---|
| Safety Invariance | Python, PyTorch, Hugging Face, CUDA | Built a one-GPU agent evaluation framework. Across 2,065 matched Qwen2.5-3B cases, aggregate security barely changed after NF4 quantization, but 73 individual security decisions flipped. |
| Epistemic Safety in AI Physics Tutors | Python, NLP, statistical analysis | Studied 60 explanations across 20 physics prompts. AI and human answers had equal median correctness, but AI explanations had lower completeness and intuition coverage. |
| PagerAgent | FastAPI, PostgreSQL, Redis, React | Built an evidence-grounded incident-response copilot that produced incident briefs in under 30 seconds and reduced manual triage and reporting effort by 70% in simulated outages. |
| OctagonRank | Node.js, Cheerio, React | Built a UFC ranking engine from 8,500+ fights, 2,500+ fighters, and 40,000+ round-level records, using 50+ features to surface trends beyond win-loss records. |
π§ͺ Research Details
Role: First author. I designed and built the evaluation pipeline for FP16, INT8, and NF4 models using native agent benchmarks.
Finding: On a controlled 65-case subset, the FP16-to-NF4 security flip rate was 36.9%, compared with a 7.7% FP16 self-repeat noise floor. Aggregate scores alone can hide meaningful behavioral changes.
Role: First author. I created the study, annotation framework, analysis pipeline, paper, and conference poster.
Finding: Human and AI explanations received the same median correctness rating, but AI explanations had lower median completeness (4.0 vs. 5.0) and intuition-bridging coverage (20% vs. 35%).
I'm interested in collaborating on AI safety evaluation, reliable agents, and research engineering.
π« abdulmohammad@berkeley.edu Β· LinkedIn
