Skip to content

Tags: harveyai/harvey-labs

Tags

v1.2.0

Toggle v1.2.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
[models] support claude 5.5 and gpt-6, send temperature only when nee…

…ded (#176)

* [models] Send temperature only to models that accept it

Claude and OpenAI adapters and judges now list the models that accept
`temperature` (a closed set of older models) instead of the models that
reject it, so a newly released model omits the parameter without a code
change. Claude models outside the older-model lists get adaptive thinking
and a 128000-token output cap by default. OpenAI-compatible servers (vLLM)
still receive `temperature` for every model.

* [models] Add Claude 5.5 and GPT-6 models to compare and the default sweep

Adds display names and list prices for Claude Opus 5.5, Opus 5, Sonnet 5.5,
Fable 5.1, GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna. The default
sweep replaces Opus 4.8 and Sonnet 5 with Opus 5.5 and Sonnet 5.5, and the
GPT-5.6 Sol/Terra/Luna tiers with GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna.

* [models] Correct list prices for Sonnet 5 and GPT-5.6 Terra and Luna

Claude Sonnet 5 is $2/$10 per 1M input/output tokens (its introductory
price became the standard price). GPT-5.6 Terra is $2/$12 and GPT-5.6
Luna is $0.20/$1.20 on OpenAI's pricing page.

* [lab-core] Remove scripts/ compatibility wrappers and refresh docs

- scripts/evaluate_submission.py and scripts/run_model_sweep.py were
  sys.path shims for the pre-package layout; the commands are
  `python -m lab_core.evaluation.run_eval` and `python -m lab_core.utils.sweep`.
- README: task-count badge (2010), a "Use As A Library" section with the
  v1.2.0 install snippet, and the command that runs the wheel's bundled
  pandoc installer, since library users do not run scripts/setup.sh.
- docs: task counts (2,010 tasks, 27 practice areas) and how commands locate
  tasks/, results/, and .env; CONTRIBUTING tree and doc-check command;
  CHANGELOG wording for the packaging entry.

---------

Co-authored-by: calvinqi <13633268+calvinqi@users.noreply.github.com>

v1.1.0

Toggle v1.1.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
[labs] add linting and typechecking (#168)

Adds linting via ruff and typechecking via pyright, generally helpful to have these in the repo for guardrails and correctness! Full CI sequence still only takes around 5min total.

## Summary

Builds on the release tooling merged in [#164](#164) by adding linting, type checks, and isolated-wheel smoke coverage.

Adds CI jobs for repo-wide Ruff, `lab_core` Pyright, and wheel import/document extraction smoke tests on Python 3.12 and 3.13. Existing Ruff violations have rule-specific file-level directives in the affected files. Existing type errors have file/rule Pyright directives, plus a single-line ignore for the optional Mistral import; runtime behavior is unchanged. Both tools are pinned in the development dependency group.

The wheel tests install only the built package and its declared runtime dependencies, then import every host module and extract known content from small synthetic XLSX/PPTX files. They run outside the checkout with Python isolation enabled and use standard-library `unittest`, so development tools and fixture generators cannot supply missing dependencies.

## Validation

- Ruff passes. Pyright basic mode passes over 30 `lab_core` source files with the recorded ignores (zero errors or warnings).
- Full offline suite: 12,790 passed, 62 skipped, 28 import subtests passed.
- Installed-wheel smoke tests and dependency consistency checks pass on Python 3.12 and 3.13.
- A new-file probe fails on missing imports, undefined names, and invalid argument types.
- The release metadata retains `openpyxl` and `python-pptx`. The isolated-wheel tests pass with those dependencies and reproduced both extraction failures when they were absent.
- The wheel-content check passes: 68 members, 146,463 bytes.

## Limits

The baseline ignores also suppress new diagnostics of the same type in the same files until removed. New files and other diagnostic types remain checked. Pyright excludes sandbox skill scripts, whose dependencies live in the container; Ruff includes them. The smoke tests cover imports and XLSX/PPTX extraction, not live provider calls or container execution.

v1.0

Toggle v1.0's commit message
Harvey LAB v1.0