Skip to content
View GBX-Max1220's full-sized avatar
🏠
Working from home
🏠
Working from home

Block or report GBX-Max1220

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
GBX-Max1220/README.md

Baixin Guo (Max)

Human-Centered AI · Human-AI Interaction · AI Evaluation

I am an undergraduate researcher in Applied Psychology working on human-AI decision calibration under epistemic uncertainty.

My research asks a simple question:

How should people rely on AI when the system's apparent confidence, precision, or reliability exceeds what its evidence can justify?

I study how AI systems communicate reliability, uncertainty, and numerical precision, and how these signals shape human reliance and decision quality.

Current Research

I am currently designing a behavioral study on how displayed numerical precision × supporting-evidence resolution affects reliance on AI-generated advice.

The core question is whether people become more willing to rely on an AI when its output appears more numerically precise—even when the available evidence does not warrant that level of precision.

Status: Study design and experimental materials in progress. No human-participant findings are claimed yet.

Research Program

My projects examine different parts of the same human-AI decision pipeline:

Project Role in the research program Evidence status
FitCalib-Bench Measuring calibration and uncertainty-expression failures in AI advice Closed engineering audit
CheckMyCoach Testing an intervention pipeline for correcting problematic AI advice Offline-evaluated prototype
Knowledge Compiler Structuring source evidence for auditable AI reasoning and evaluation Evidence infrastructure
InteractionKit Experimental infrastructure for structured Human-AI interaction studies Frozen methodological asset · pending human validation
CalTrust Exploring adaptive assistance based on observed trust and reliance signals Frozen simulation-only prototype

Together, these projects explore a broader problem:

measurement → intervention → evidence → behavioral evaluation → adaptive assistance

Research Interests

  • Human-AI Interaction
  • Human-AI Decision Making
  • AI Evaluation & Calibration
  • Uncertainty and Confidence Communication
  • Trust, Reliance, and Overreliance
  • Interactive / Human-Centered AI

Evidence Discipline

I try to keep a strict boundary between different forms of evidence.

Software implementation, simulation, automated evaluation, and human behavioral evidence are not interchangeable.

Public project pages therefore distinguish engineering evidence from scientific evidence, and I do not claim human-participant findings where human studies have not yet been conducted.

Links

Pinned Loading

  1. CalTrust CalTrust Public

    Adaptive Trust Calibration via Contextual Bandits

    Python

  2. CheckMyCoach CheckMyCoach Public

    A Calibration Pipeline for LLM Uncertainty Expression

    Python

  3. FitCalib-Bench FitCalib-Bench Public

    Benchmarking LLM calibration in safety-critical advice domains

    Python

  4. InteractionKit InteractionKit Public

    TypeScript

  5. knowledge-compiler knowledge-compiler Public

    Structured exercise-science evidence infrastructure with provenance auditing and recovery

    Python

  6. maxguo.dev maxguo.dev Public

    personal website

    Astro