Skip to content
View melcheikh's full-sized avatar
🏠
Working from home
🏠
Working from home

Block or report melcheikh

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
melcheikh/README.md

Martín el Cheikh

AI engineer working on model evaluation and behavior. I come from psychology (UBA), and that is the lens I bring to how models are measured, not just whether they run.

What I'm working on

  • Evals. I send measurement-validity fixes to inspect_evals (UK AI Safety Institute): a grader failure counted as a real score instead of a missing value, a verdict parsed wrong, quietly biasing what the eval reports. My PRs.

  • Inference. I send fixes to vLLM. I run models on consumer Blackwell (RTX 5090, SM120, NVFP4) and chase failures the project's CI does not cover. My PRs.

  • Models. NVFP4 quantizations of Gemma 4 on Hugging Face.

The angle

Most evals are written without experimental design: no control condition, no construct validity, no measure of uncertainty. A judge model that refuses the worst transcripts, or a scorer that reports zero when it measured nothing, makes a model look safer than it is. That gap, between "the number moved" and "the number means something", is where I work.

Background

Psychology (UBA, in progress). Autodidact in AI since 2024. Buenos Aires.

Pinned Loading

  1. inspect_evals inspect_evals Public

    Forked from UKGovernmentBEIS/inspect_evals

    Collection of evals for Inspect AI

    Python

  2. vllm vllm Public

    Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python

  3. AudioLegion AudioLegion Public

    Linux 4.0 Audio & Volume Scaling Fix for Lenovo Legion Pro 7i Gen 10 (16IAX10H / AW88399)

    Shell 1

  4. Flux2 Flux2 Public

    Python 1