Surge AI’s cover photo
Surge AI

Surge AI

Software Development

San Francisco, California 34,735 followers

Human intelligence for AGI

About us

Our mission is to raise AGI with the richness of humanity — curious, witty, imaginative, and full of breathtaking brilliance.

Website
https://www.surgehq.ai
Industry
Software Development
Company size
51-200 employees
Headquarters
San Francisco, California
Type
Privately Held
Founded
2020
Specialties
machine learning, data labeling, artificial intelligence, and software

Products

Locations

Employees at Surge AI

Updates

  • For their Mistral Large 4 release, Mistral ran a blind Surge human eval, with our expert software engineers evaluating five frontier models on coding quality. Le Chonk ranked #1 among open-weight models and #2 overall, behind only Opus 5. Unit tests ask: does it work? Human code review asks: would you merge it? Our evals measures the difference. Working code is the floor; professional taste and judgment are what make it shippable. Congrats to the Mistral team on le Chonk and their big release 🇫🇷

    View organization page for Mistral

    728,896 followers

    Introducing Mistral Large 4, aka Le Chonk. Mistral Large 4 is the best open model for sovereign AI, with state-of-the-art, open-weight performance in cyber defense, manufacturing, finance and multimodal tasks, and a solid foundation for customizing models across key industries. It is a 1 trillion-parameter multimodal model with 49 billion active parameters, trained from scratch on close to 4,000 Grace Blackwell GPUs and deployed in our own data centers. Its performance is competitive with the strongest open models globally, significantly outperforming any open-weight model developed in the US or Europe, despite using a fraction of the GPU training resources. Mistral Large 4 is available to preview today with full release later this month—your feedback will help shape what ships next.

  • Surge AI reposted this

    🥖 LGTM, le Chonk 🥖 in their latest Mistral Large 4 release, Mistral ran a blind Surge human eval across five frontier models - professional software engineers comparing five frontier models on coding quality. the big guy finished #2, first among open-weight models and behind only Opus 5. most benchmarks can narrowly measure correctness. but human evals get at the deeper stuff: the professional taste, judgment, and maintainability we look for in the real world as well the difference between passing a set of unit tests vs. getting a "LGTM - ship it" bien joué, le Chonk ! congrats to the entire Mistral team 🇫🇷

    • No alternative text description for this image
  • Surge AI reposted this

    A big focus for our team at Pulse is making the data AI agents receive more useful. Huge congrats to the Surge AI team on GDP.xlsx, a strong benchmark testing spreadsheet reasoning across finance, engineering and healthcare. With Ultra 2, our latest model, we’re focused on preserving hidden content, formulas, complete records and visual context within these enterprise documents. We evaluated Pulse’s spreadsheet extraction pipeline using the benchmark’s OpenHands agent harness, providing structured text and image evidence upfront. Our latest reported task-pass rates, compared with our earlier Pulse-assisted results: - Claude Opus 5.5 46.4% (up from 41.4%; +5 percentage points) - GPT-6.1 Sol 37.9% (up from 32.9%; +5 percentage points) - GPT-6 Astra 35.0% (up from 30.0%; +5 percentage points) - Claude Fable 5.1 33.6% (up from 28.6%; +5 percentage points) - Gemini 3.8 Flash 23.6% (up from 18.6%; +5 percentage points) Better data gives agents more to work with, read more in the link below.

    • No alternative text description for this image
  • Surge AI reposted this

    GDP.pdf just got a little brother: introducing GDP.xlsx the world runs on spreadsheets. somewhere in Sheet2!D18 there's a vlookup that changes a $14m decision. GDP.xlsx has 70 real-world tasks across 12 domains. what's hard is figuring out how the workbook even works: → which tab is stale? → what does a red cell mean? → which comment explains the exceptions? → which chart does Finance actually trust? → did you remember to check the termination clauses in Sheet 3? the best frontier agent scores 38.3%. (congrats Gemini 4 Argon!) blog post: surgehq.ai/blog/gdp-xlsx

    • No alternative text description for this image

Similar pages

Browse jobs