Agent evals under uncertainty.
Mohamed A M Elansary, PhD — candidate for Member of Technical Staff - Research. Measurement discipline for agent simulation, scalable oversight, and frontier evaluation — with honest gaps on RL-environment depth and San Francisco five-day onsite relocation.
Scientific evaluation
- Designed multimodel, multi-basin forecast experiments across hydroclimates.
- Quantified and reduced uncertainty using imperfect USGS, NOAA, and NASA observations.
- Ran reproducible Linux/HPC workflows and Python, R, and Bash data pipelines.
Production delivery
- Builds agentic LLM workflows using GPT, Claude, and Gemini.
- Maintains regression evaluation sets for multi-tenant agent workflows.
- Ships retrieval, routing, isolation, provenance, validation, and monitoring systems.
What I would demonstrate
Define a small agent task suite, record configuration, inspect tool-using trajectories, separate capability failures from tooling or metric gaming, quantify uncertainty, and deliver a decision-ready memo grounded in the evidence.
Honest fit boundary
San Francisco five-days-a-week onsite is a household relocation-with-package ask; I do not assert remote or hybrid. RL-environment and SWE-agent simulation depth is not sourced. I have not claimed GRPO/SFT ownership, RLHF, or invented safety research or metrics. My research record is environmental engineering and hydrologic forecasting; the transfer is multimodel evaluation under uncertainty and production agent evaluation sets.
Role and logistics
San Francisco 5 days/week, JD quote (re-verified 2026-09-03): “To support close collaboration, this role is based in our San Francisco headquarters and requires in-office attendance five days a week.”
Compensation, JD quote: “The expected base salary range for this role is $175,000 - $300,000 USD. In addition to base salary, we offer equity and benefits.”
Clearance / citizenship / ITAR: none found in the live JD or Greenhouse job content (re-verified 2026-09-03).
Official role: Patronus AI / Greenhouse