AI engineer working on model evaluation and behavior. I come from psychology (UBA), and that is the lens I bring to how models are measured, not just whether they run.
-
Evals. I send measurement-validity fixes to inspect_evals (UK AI Safety Institute): a grader failure counted as a real score instead of a missing value, a verdict parsed wrong, quietly biasing what the eval reports. My PRs.
-
Inference. I send fixes to vLLM. I run models on consumer Blackwell (RTX 5090, SM120, NVFP4) and chase failures the project's CI does not cover. My PRs.
-
Models. NVFP4 quantizations of Gemma 4 on Hugging Face.
Most evals are written without experimental design: no control condition, no construct validity, no measure of uncertainty. A judge model that refuses the worst transcripts, or a scorer that reports zero when it measured nothing, makes a model look safer than it is. That gap, between "the number moved" and "the number means something", is where I work.
Psychology (UBA, in progress). Autodidact in AI since 2024. Buenos Aires.