Only 2 of 66 runs caught a $25.7M pricing error. One screenshot. One unit mistake. A 100× overstatement. See what else frontier models miss on DAYJOB: Finance: https://lnkd.in/efWEU-Gz
Surge AI
Software Development
San Francisco, California 34,735 followers
Human intelligence for AGI
About us
Our mission is to raise AGI with the richness of humanity — curious, witty, imaginative, and full of breathtaking brilliance.
- Website
-
https://www.surgehq.ai
External link for Surge AI
- Industry
- Software Development
- Company size
- 51-200 employees
- Headquarters
- San Francisco, California
- Type
- Privately Held
- Founded
- 2020
- Specialties
- machine learning, data labeling, artificial intelligence, and software
Products
Surge AI
Data Labeling Platforms
Surge AI provides both the platform and the workforce for high-skill data labeling, including NLP, code generation, search evaluation, adversarial training, and much more.
Locations
-
Primary
Get directions
San Francisco, California 94105, US
-
Get directions
New York City, US
Employees at Surge AI
Updates
-
For their Mistral Large 4 release, Mistral ran a blind Surge human eval, with our expert software engineers evaluating five frontier models on coding quality. Le Chonk ranked #1 among open-weight models and #2 overall, behind only Opus 5. Unit tests ask: does it work? Human code review asks: would you merge it? Our evals measures the difference. Working code is the floor; professional taste and judgment are what make it shippable. Congrats to the Mistral team on le Chonk and their big release 🇫🇷
Introducing Mistral Large 4, aka Le Chonk. Mistral Large 4 is the best open model for sovereign AI, with state-of-the-art, open-weight performance in cyber defense, manufacturing, finance and multimodal tasks, and a solid foundation for customizing models across key industries. It is a 1 trillion-parameter multimodal model with 49 billion active parameters, trained from scratch on close to 4,000 Grace Blackwell GPUs and deployed in our own data centers. Its performance is competitive with the strongest open models globally, significantly outperforming any open-weight model developed in the US or Europe, despite using a fraction of the GPU training resources. Mistral Large 4 is available to preview today with full release later this month—your feedback will help shape what ships next.
-
Surge AI reposted this
🥖 LGTM, le Chonk 🥖 in their latest Mistral Large 4 release, Mistral ran a blind Surge human eval across five frontier models - professional software engineers comparing five frontier models on coding quality. the big guy finished #2, first among open-weight models and behind only Opus 5. most benchmarks can narrowly measure correctness. but human evals get at the deeper stuff: the professional taste, judgment, and maintainability we look for in the real world as well the difference between passing a set of unit tests vs. getting a "LGTM - ship it" bien joué, le Chonk ! congrats to the entire Mistral team 🇫🇷
-
-
Surge AI reposted this
A big focus for our team at Pulse is making the data AI agents receive more useful. Huge congrats to the Surge AI team on GDP.xlsx, a strong benchmark testing spreadsheet reasoning across finance, engineering and healthcare. With Ultra 2, our latest model, we’re focused on preserving hidden content, formulas, complete records and visual context within these enterprise documents. We evaluated Pulse’s spreadsheet extraction pipeline using the benchmark’s OpenHands agent harness, providing structured text and image evidence upfront. Our latest reported task-pass rates, compared with our earlier Pulse-assisted results: - Claude Opus 5.5 46.4% (up from 41.4%; +5 percentage points) - GPT-6.1 Sol 37.9% (up from 32.9%; +5 percentage points) - GPT-6 Astra 35.0% (up from 30.0%; +5 percentage points) - Claude Fable 5.1 33.6% (up from 28.6%; +5 percentage points) - Gemini 3.8 Flash 23.6% (up from 18.6%; +5 percentage points) Better data gives agents more to work with, read more in the link below.
-
-
Surge AI reposted this
they grow up so fast 😊 #1 Riemann-bench #1 Chartography #1 GDP.xlsx congrats, Argon and Google DeepMind! this one’s going on the fridge.
-
-
Surge AI reposted this
GDP.pdf just got a little brother: introducing GDP.xlsx the world runs on spreadsheets. somewhere in Sheet2!D18 there's a vlookup that changes a $14m decision. GDP.xlsx has 70 real-world tasks across 12 domains. what's hard is figuring out how the workbook even works: → which tab is stale? → what does a red cell mean? → which comment explains the exceptions? → which chart does Finance actually trust? → did you remember to check the termination clauses in Sheet 3? the best frontier agent scores 38.3%. (congrats Gemini 4 Argon!) blog post: surgehq.ai/blog/gdp-xlsx
-
-
Surge AI reposted this
other models: "the tenant's liability is unlimited" gpt-6.1: "not according to page 14" great job, sol
-
-
A benchmark shouldn’t just be hard to climb. Climbing it should make the model stronger in ways that transfer. So what does GDP.pdf teach, beyond PDFs? Turns out, quite a lot. New research: https://lnkd.in/gSQXZ-8s Paper: https://lnkd.in/gGJkBZwU