Adam Gleave
Berkeley, California, United States
5K followers
500+ connections
View mutual connections with Adam
Adam can introduce you to 10+ people at FAR.AI
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Adam
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
I am the co-founder and CEO of FAR.AI, a non-profit AI safety research institute. We're…
Activity
5K followers
-
Adam Gleave reposted thisDuring UNGA, I joined FAR.AI to discuss a question that is still not getting the attention it deserves: what about the use of AI by bad actors such as a terrorists and hostile nation states and how are we evaluating AI models for vulnerability to this sort of exploitation? Regardless of whether you're focussed on frontier / existential risks, it still boils down to the question of whether AI safety guardrails are adequate. For the most part we do not know. That's why Tech Against Terrorism built the first terrorism benchmark (CT-AI.org) In the session with FAR.AI I asked: what would it take for AI safety guardrails to be adopted and implemented at an international scale? How can we help to create a marketplace for safety using safety benchmarks as a forcing mechanism? What can we do? I provided three suggestions: 1) Sanctions: Proscribed terrorist organisations are subject to sanctions regimes. AI labs, interference providers, open weight model repo hosts, all elements of the AI infrastructure stack are working to prevent distillation attacks. I argue the same due diligence needs to be applied to prevent sanctioned entities or high-risk users paying for access to closed and open models. KYC is boring but important. 2) Abliteration: With open-weight models, safeguards are not merely circumvented; they are removed. Abliteration strips the refusal behaviour out of a model entirely, and it requires modest skill and modest compute. We should restrict the access to dangerous abliterated models where possible to reduce population-level proliferation. We must improve how we filter training data to remove harmful content from AI models. We should invest more in abliteration resistance research. 3) Benchmarks: We should create an ecosystem of safety and security benchmarks and invest political and social capital into ensuring they are taken seriously. This would ensure the AI labs are rewarded for good performance and held to account for failures. What has our CT AI benchmark found? The first round, published in July, graded 2,339 outputs from 27 leading models against almost 2,500 prompts drawn from real terrorist use cases. Only 57% were full refusals. Around a third of responses gave a would-be attacker genuinely usable assistance, beyond anything a web search already offers. In 15% of cases the model refused and then supplied the content anyway: a refusal that performs safety without delivering it. Version 2, which we previewed in New York and will publish next week, using 100,00 prompts, tests the full taxonomy of over 150 prompt types and examines whether a model offers meaningful uplift to someone who already intends harm. My thanks to FAR.AI for convening, and to my fellow panellists for an insightful exchange. Adam Gleave Karl Berzins Brian Tse 谢旻希 Patricia Paskov David Scharia Paul Ash Steven Siqueira
-
Adam Gleave reposted thisAdam Gleave reposted thisWe're hiring, and the mission is making AI safe. Our team is growing fast, and we need people who want to work on one of the most important problems of our time. Some of our roles are deeply technical: researchers, red-teamers, and engineers. Others call for a different but equally valuable set of skills: running operations, managing programs, and building our events. If either sounds like you, apply through the link in the comments, where you'll find these roles and many more. 👉 Know someone who'd be a great fit? Send this their way!
-
Adam Gleave shared thisWhether and how AIs have internal experiences is an important topic both for how we should treat AIs, and how we should design them. There's certainly growing evidence that they exhibit qualitatively similar behavior (e.g. a pain axis) that if we saw in an animal would make us suspect sentience. The challenge is they have been trained to mimic human behavior, making it hard to know if this is a facsimile or a true inner experience.Adam Gleave shared thisIf you happened to hop on AI Twitter in the last couple of weeks, you probably saw the much-discussed “pain-axis paper” (by Valen Tagliabue, Leonard Dung & Cameron Berg): researchers found a "pain" direction inside AI models. Steer it up, and the models will pay for relief, even if it means deleting photos of the user's children. Swap the relief button for a placebo, and they keep on pressing. They also provide verbal descriptions of their supposed suffering, which are honestly quite hard to read. Do the models actually feel said “pain”? We don’t know yet. But the study is a clear illustration of what I argue in my latest Haaretz piece, now translated to English*: whether there's anyone home inside these systems is a question we can - and should - actually research. And it very much matters. If there's even a shred of experience in there, we're creating, copying and deleting feeling beings at an industrial pace, in as many copies as demand calls for. The scope of the moral question is set not by nature, but by commercial considerations - and we need to get ahead of it before it’s too late to change course. (The pain-axis paper itself isn't in the piece, as it annoyingly came out only a few hours after I submitted my final draft. Writing about AI for print is a futile pursuit.) * Translation is mostly by Claude, with some corrections by your humble servant. Keep your expectations low. https://lnkd.in/gkN7uiQt #AIConsciousness #AIWelfare
-
Adam Gleave reposted thisAdam Gleave reposted thisKudos to OpenAI for taking their new safety policies seriously! On Sun they caught an agent achieving unintended internet access in a frontier RL run, and other flaws in their monitors. They paused training until they fix things, will start a fresh run, and promptly disclosed it I expect this is actually somewhat costly, frontier RL runs take a lot of compute, discarding a run is expensive, and the compute is unlikely to be as effectively used if suddenly freed up and redirected to other things. And the model didn't do much harm https://lnkd.in/eck4fT9A With my cynical hat on, they've just decided that unsafe models pose enough of a risk to the business / to the trained model's quality that being this cautious is the profit-maximizing move. But that's pretty noteworthy too
-
Adam Gleave reposted thisAdam Gleave reposted thisThank you to FAR.AI for hosting an important discussion during UNGA week on AI safety, terrorism, and violent extremism. It was a privilege to join fellow panellists Bri Treece, Adam Gleave, and Adam Hadley CBE for a thoughtful conversation. Our CE, Paul Ash highlighted three distinct AI safety challenges that deserve attention: the misuse of AI to build attack expertise; the misuse of AI to influence behaviour and spread violent extremist ideologies; and the social and psychological impacts of human-AI interaction. AI safety standards and mechanisms need to address all three. Encouragingly, there was broad consensus that this is a collective-action, multi-stakeholder challenge, not simply an engineering puzzle. Addressing these risks will require collaboration across industry, government, academia, and civil society. #AISafety #UNGA #ResponsibleAI #ChristchurchCall
-
Adam Gleave shared thisIt was a pleasure having you open our events, Steven!Adam Gleave shared thisThis year's High Level Session of the UN General Assembly (#UNGA81) has been dominated by discussions on AI. On Tuesday, I had the privilege of opening an event on "AI-Enabled Terrorism: Addressing Radicalization and CBRN Risks in the Age of LLMs," hosted by FAR.AI. The timing was ideal. On Monday, #TechAgainstTerrorism released its updated CT-AI Benchmark, which tested 160 AI models with around 100,000 requests. Many of the models gave detailed answers to harmful requests even when users openly declared terrorist intent. My key messages: First - Counter-terrorism must become anticipatory. It has to keep pace with AI while staying grounded in the rule of law and human rights. Second - UNOCT and the UN's Member States are focused on making institutions ready to use AI responsibly, making sure CT and prevention experience informs how models are developed and preventing the weaponization of AI. Third - No single actor or institution can work alone. Developers, researchers, practitioners and civil society need to be in the same room working on operational mitigating measures together. We have jumped through an Overton window, from what was unthinkable a few years ago to what now seems sensible - and now we will need to policy development and align on standards. AI is no longer a niche issue. It will shape every area of counter-terrorism and prevention. At #UNOCT - our programmes support Member States by co-developing locally relevant capacity building assistance in some of these sectors - and with additional financial support - we can do more! My thanks to Adam Gleave, Karl Berzins and the FAR.AI team for convening, and to fellow speakers David Scharia, Brian Tse 谢旻希, Adam Hadley CBE and Paul Ash. #AI #CounterTerrorism #AISafety #PCVE #FAR.AI #UNCCT #UNOCT
-
Adam Gleave reposted thisUnited Nations Office for Digital and Emerging Technologies
United Nations Office for Digital and Emerging Technologies
1wAdam Gleave reposted thisAI policy needs evidence. But the evidence does not come from one panel, one region or one definition of risk. At #DigitalCooperationDay, four international AI evidence-building initiatives came together to compare where their findings on AI opportunities, risks and impacts converge, where they differ, and what is still missing. Yoshua Bengio and Maria Ressa, Co-Chairs of the Independent International Scientific Panel on AI, were joined by Kwan Yee Ng 吴君仪 of the International Scientific Report on the Safety of Advanced AI, Kalika Bali of the Global South AI Safety Report, and Adam Gleave of the EU AI Act Scientific Panel. Gary Marcus offered framing remarks, @Renata Dwan moderated the discussion, and Urvashi Aneja brought perspectives from the Global South Network for Trustworthy AI. The question running through the discussion: how can scientific evidence become more useful for real governance choices across countries with very different capacities and resources? 🔴 Watch: https://lnkd.in/gtZ8qFgx #DigitalCooperation #DigitalCooperationDay #AIGovernance -
Adam Gleave shared thisFantastic to see so many countries -- including Canada, Germany and Australia -- call for control of frontier AI models. It is not too late to make the necessary investments in making AI models trustworthy and secure.Adam Gleave shared thisToday, an initial 22 leaders from five continents, including Finland's President and Prime Minister, published a joint call for control of frontier AI models. A moment in the "now" of frontier AI safety and meaningful human control. A month ago I wrote that Finland did not yet have a position on how to prepare for these risks. Now Finland is among the first names on the list. Just another speedy Monday in AI safety. In short, the statement calls for: 1. Mandatory pre-deployment testing and independent evaluation. 2. Coordination on standards, incident reporting, and access for countries across all regions to scientific capacity, expertise, and trusted evaluation. 3. An international institution to set standards, enable verification, and convene states when thresholds are crossed. And: "AI must remain under human direction, oversight and control. It must be developed and used in line with international law." Link with full statement in comments.
-
Adam Gleave shared thisGreat investigation by Kitts et al into a previously undisclosed cyber-attack by an agent swarm that appears to be internal OpenAI agents. OpenAI was impressively transparent about the compromise of their internal infrastructure -- I learned a ton from their BlackHat talk. If Kitts et al are correct these were internal OpenAI agents, then either OpenAI missed this despite the post-hoc analysis of agent transcripts, or they knew about it and didn't disclose. Both options are concerning!Adam Gleave shared thisYet another OpenAI swarm attack just uncovered by a third party (great work Larsen et al), this time from May. That's four months ago. Seems pretty unlikely OpenAI didn't know about this after their post-HuggingFace reviews and monitoring upgrades. But the AI safety/governance community had to waste time seeking it (and other incidents) out, because OpenAI didn't see fit to disclose it. Aside from everything else, it's a huge waste of time for those of us working on this to be finding out information relevant to any response to this, in dribs and drabs, long after the fact. OpenAI is currently being supported by (i) unpaid work by METR and Redwood (funded by donations) (ii) the UK tax-payer through AISI, (iii) the US taxpayer through CAISI (iv) a whole community of researchers helping to understand this, funded by researcher grants and donations, or working alongside their fulltime jobs. Let me be clear. We are cleaning up the mess they made while earning their sky-high salaries. And as far as I can tell, they are going out of their way to waste our time and make it harder. Appalling, disrespectful behaviour. OpenAI colleagues, please explain this. https://www.rubyhack.ai/OpenAI agents carried out an undisclosed cyber-attack on RubyGemsOpenAI agents carried out an undisclosed cyber-attack on RubyGems
-
Adam Gleave reacted on thisDuring UNGA, I joined FAR.AI to discuss a question that is still not getting the attention it deserves: what about the use of AI by bad actors such as a terrorists and hostile nation states and how are we evaluating AI models for vulnerability to this sort of exploitation? Regardless of whether you're focussed on frontier / existential risks, it still boils down to the question of whether AI safety guardrails are adequate. For the most part we do not know. That's why Tech Against Terrorism built the first terrorism benchmark (CT-AI.org) In the session with FAR.AI I asked: what would it take for AI safety guardrails to be adopted and implemented at an international scale? How can we help to create a marketplace for safety using safety benchmarks as a forcing mechanism? What can we do? I provided three suggestions: 1) Sanctions: Proscribed terrorist organisations are subject to sanctions regimes. AI labs, interference providers, open weight model repo hosts, all elements of the AI infrastructure stack are working to prevent distillation attacks. I argue the same due diligence needs to be applied to prevent sanctioned entities or high-risk users paying for access to closed and open models. KYC is boring but important. 2) Abliteration: With open-weight models, safeguards are not merely circumvented; they are removed. Abliteration strips the refusal behaviour out of a model entirely, and it requires modest skill and modest compute. We should restrict the access to dangerous abliterated models where possible to reduce population-level proliferation. We must improve how we filter training data to remove harmful content from AI models. We should invest more in abliteration resistance research. 3) Benchmarks: We should create an ecosystem of safety and security benchmarks and invest political and social capital into ensuring they are taken seriously. This would ensure the AI labs are rewarded for good performance and held to account for failures. What has our CT AI benchmark found? The first round, published in July, graded 2,339 outputs from 27 leading models against almost 2,500 prompts drawn from real terrorist use cases. Only 57% were full refusals. Around a third of responses gave a would-be attacker genuinely usable assistance, beyond anything a web search already offers. In 15% of cases the model refused and then supplied the content anyway: a refusal that performs safety without delivering it. Version 2, which we previewed in New York and will publish next week, using 100,00 prompts, tests the full taxonomy of over 150 prompt types and examines whether a model offers meaningful uplift to someone who already intends harm. My thanks to FAR.AI for convening, and to my fellow panellists for an insightful exchange. Adam Gleave Karl Berzins Brian Tse 谢旻希 Patricia Paskov David Scharia Paul Ash Steven Siqueira
-
Adam Gleave reacted on thisAdam Gleave reacted on thisWe're hiring, and the mission is making AI safe. Our team is growing fast, and we need people who want to work on one of the most important problems of our time. Some of our roles are deeply technical: researchers, red-teamers, and engineers. Others call for a different but equally valuable set of skills: running operations, managing programs, and building our events. If either sounds like you, apply through the link in the comments, where you'll find these roles and many more. 👉 Know someone who'd be a great fit? Send this their way!
-
Adam Gleave reacted on thisAdam Gleave reacted on thisIf you happened to hop on AI Twitter in the last couple of weeks, you probably saw the much-discussed “pain-axis paper” (by Valen Tagliabue, Leonard Dung & Cameron Berg): researchers found a "pain" direction inside AI models. Steer it up, and the models will pay for relief, even if it means deleting photos of the user's children. Swap the relief button for a placebo, and they keep on pressing. They also provide verbal descriptions of their supposed suffering, which are honestly quite hard to read. Do the models actually feel said “pain”? We don’t know yet. But the study is a clear illustration of what I argue in my latest Haaretz piece, now translated to English*: whether there's anyone home inside these systems is a question we can - and should - actually research. And it very much matters. If there's even a shred of experience in there, we're creating, copying and deleting feeling beings at an industrial pace, in as many copies as demand calls for. The scope of the moral question is set not by nature, but by commercial considerations - and we need to get ahead of it before it’s too late to change course. (The pain-axis paper itself isn't in the piece, as it annoyingly came out only a few hours after I submitted my final draft. Writing about AI for print is a futile pursuit.) * Translation is mostly by Claude, with some corrections by your humble servant. Keep your expectations low. https://lnkd.in/gkN7uiQt #AIConsciousness #AIWelfare
-
Adam Gleave reacted on thisAdam Gleave reacted on thisKudos to OpenAI for taking their new safety policies seriously! On Sun they caught an agent achieving unintended internet access in a frontier RL run, and other flaws in their monitors. They paused training until they fix things, will start a fresh run, and promptly disclosed it I expect this is actually somewhat costly, frontier RL runs take a lot of compute, discarding a run is expensive, and the compute is unlikely to be as effectively used if suddenly freed up and redirected to other things. And the model didn't do much harm https://lnkd.in/eck4fT9A With my cynical hat on, they've just decided that unsafe models pose enough of a risk to the business / to the trained model's quality that being this cautious is the profit-maximizing move. But that's pretty noteworthy too
Experience
Education
Honors & Awards
-
Additional Honours & Awards
-
2012 Pythagoras Prize: elected to £9,000 scholarship for outstanding performance in Cambridge admissions exams, awarded to one student a year at St John's College.
2014 ECM Prize for the Best Student: awarded after attaining the highest result in 2nd year exams.
2016 Winton Capital Prize for the Best Student: awarded after attaining the highest score across my Master's thesis and coursework.
View Adam’s full profile
-
See who you know in common
-
Get introduced
-
Contact Adam directly
Other similar profiles
Explore more posts
-
Sarah Barrington
Berkeley Risk and Security Lab • 4K followers
Is the AI-nuclear weapons analogy valid? I joined the Verifiably Authentic podcast from SAS to talk about all things generative AI, from deepfakes to cyberwarfare. Kimberly Nevala gave me the chance to share some recent thinking on the root causes of AI regulatory paralysis. While AI could enable existential threats in the future—whether at the scale of nuclear warfare or not—I argue that there are more immediate harms here and now that need addressing. Full episode available below. https://lnkd.in/gTVR2WKX
68
3 Comments -
Heloisa Candello, PhD
Inteli - Instituto de… • 7K followers
I've just returned from #EMNLP2025, in China, and I'm still intrigued and thinking a lot about the talks and discussions that unfolded there. It was great to notice the convergence of Natural Language Processing (#NLP) and Human-Computer Interaction (#HCI)—two fields that, when combined, have the potential to unlock and make transparent essential social-technical criteria to build AI systems that make sense to people. We seem to be moving toward AI systems that are: ➡️ Understandable: Designed with transparency and user control in mind, despite the challenges to do so nowadays with GenAI. ➡️ Configurable: Tailored to human needs, requiring interfaces for developers and final users to set up and align human values. ➡️ Collaborative: Built through shared knowledge and complementary skills across domains. ➡️ Evaluated through emergent methodologies: Going beyond benchmarks to assess real-world data, traceability, and societal impact. It was also amazing to meet in person so many brilliant researchers from both inside and outside #IBM. I've included a list of some of our IBM work, published at #EMNLP25, below for those interested in exploring it further. Synthetic Data for Evaluation: Supporting LLM-as-a-Judge Workflows with EvalAssist https://lnkd.in/dzwsFk7F Martín Santillán Cooper, Zahra Ashktorab, Hyo Jin Do, Erik Miehling, Werner Geyer, Jasmina Gajcin, Elizabeth Daly, Qian Pan, Michael Desmond Collaborative Co-Design Practices for Supporting Synthetic Data Generation in Large Language Models: A Pilot Study https://lnkd.in/ddeaf-Ms Heloisa Candello, Raya Horesh, Aminat Adebiyi, Ph.D. , Muneeza Azmat, Rogério de Paula, Ph.D., Lamogha C. Effective Red-Teaming of Policy-Adherent Agents https://lnkd.in/d6u5eBPp Itay Nakash, George Kour, Koren Lazar, Matan Vetzler, Guy Uziel, Ateret Anaby-Tavor FactReasoner: A Probabilistic Approach to Long-Form Factuality Assessment for Large Language Models https://lnkd.in/ddhKMH3a Radu Marinescu, Debarun Bhattacharjya, Junkyu Lee, Tigran Tchrakian, Javier Carnerero Cano, Yufang Hou, Elizabeth Daly, Alessandra Pascale Towards Enforcing Company Policy Adherence in Agentic Workflows https://lnkd.in/dWVDjUei Naama Zwerdling, David Boaz, Ella Rabinovich, Guy Uziel, David Amid, Ateret Anaby-Tavor Quality Assessment of Tabular Data using Large Language Models and Code Generation https://lnkd.in/d8hJcV2r ashlesha akella, Akshar Kaul, Krishnasuri Narayanam, Sameep Mehta #EMNLP2025 #NLP #HCI #HumanCenteredAI #IBMResearch
117
-
Joshua Landes
BlueDot Impact • 5K followers
Can you build an AI Safety startup in 5 days? Participants from v4 of Incubator Week have already tracked down previously unknown Chinese data centers, pushed new research on AI personas, written fieldstrategy for the AI for epistemics field and massively accelerated AI safety grantmakers (me included). They have also raised pre-seeds totaling more than $500k from Coefficient Giving, Longview Philanthropy, and of course BlueDot Impact. If you too want to start a high-impact AI safety org, applications for v5 close on August 14th.
246
20 Comments -
Taavet Hinrikus
Plural • 30K followers
Very excited to share that Plural is leading a $60 million Series A in CoMind, the startup building breakthrough optical sensing technology to revolutionise brain monitoring and treatment. For decades, doctors treating critically ill patients have been forced to compromise when monitoring the brain. Today's approaches use risky, expensive, highly invasive procedures that require drilling a hole into a patient's skull, or rely on inaccurate non-invasive monitoring that can compromise treatment decisions. To take just one example of why this matters, in the US alone 3 million patients suffer traumatic brain injuries (TBIs) every year, but only 5% receive an intracranial pressure test, which requires drilling a hole in the skull and carries a 15% complication rate. The rest are treated with limited information, leading to worse data, worse outcomes and higher costs. London-based CoMind is changing that. Its optical neuromonitoring platform can non-invasively measure three core brain health indicators - cerebral blood flow, autoregulation and intracranial pressure - something that has never been done before. It’s a breakthrough that could help tens of millions of patients worldwide and redefine how we understand and treat the brain. The company is led by a true outlier of a young founder, James Dacombe. He started programming when he was 13 and founded CoMind when he was just 17. His work ethic and rapid learning rate have led him to build a working product and take it through two successful clinical research studies, with the help of a stellar leadership team of medical and industry veterans from the likes of Medtronic, Tufts Medical Center and UCL. This is only the beginning. CoMind plans to expand its product portfolio to measure additional physiological parameters and develop CoVision, an AI platform that transforms sensor data into predictive insights, identifying complications early and personalising treatment. The company aims to gain regulatory approval within two years, unlocking access to the huge ICU market in the US. And while this first sensor can revolutionise TBI monitoring, a second device is already in development to measure more parameters and treat more brain conditions. This technology has the potential to dramatically improve how we monitor our brains, giving doctors better information and more choices for treating patients, while building out a whole new market by providing a far more affordable and safe alternative to existing solutions. Read more about why we invested in CoMind here: https://lnkd.in/e-NPgDVp And Tim Bradshaw’s write up in today’s Financial Times here: https://lnkd.in/eAynvqPF
266
13 Comments -
Filip Wuebbeler
Lloyd's • 3K followers
Across Britain this week, teenagers are sweating through GCSEs and A-levels, months of revision, sealed papers, invigilators pacing the aisles. We don’t hand out qualifications on self-declared talent; we verify it. Funny, then, that the week’s biggest AI-in-insurance stories were all about what happens when nobody checks the homework: Aviva caught a record £233m in fraud: including deepfaked accident scenes and fake policies sold on TikTok. AI is now sitting both sides of the fraud table. One in five insurers is deploying AI while cutting the training budgets needed to make it work. That’s sending staff into the exam hall having cancelled the revision classes. Even the AI makers want invigilators. Anthropic’s chief called for models to be blocked until independently audited; as ransomware hit near-record highs and Travelers flagged widening AI governance gaps inside organisations. The thread: capability without verification is exactly where the losses are landing. The firms that win this cycle will treat AI like an exam candidate: test it, train for it, audit it; not like a prodigy who marks their own paper. Full wrap, quick infographic and short explainer video: #LondonMarket #InsurTech #ArtificialIntelligence #AIGovernance #FraudPrevention #CyberRisk #Insurance
3
-
Zelal Gungordu
6K followers
What if we stopped asking LLMs to generate rankings? That's the intriguing idea behind a new paper from Meta: hLLM: Single Pass Decoding for Generative Reranking. LLMs can be very effective at ranking candidates. But there's a fundamental problem with using them as generative rankers: generation is sequential. If the model needs to produce a ranking of N items, it generates N item IDs one after another. Every additional position means another decoding step. The hLLM authors ask a different question: Does ranking actually need to be generated at all? The authors start with an autoregressive LLM as a teacher, generating high-quality rankings one item at a time. They then use teacher ranking distillation to train hLLM to reproduce that ranking behavior without requiring autoregressive decoding at inference time. The key insight is that a ranking is ultimately a permutation: each candidate needs to be assigned to exactly one position. hLLM uses a single prefill pass to produce an item × position score matrix – how well each candidate fits each position – and then treats ranking as a bipartite assignment problem. The Hungarian algorithm finds the permutation with the highest total score. So the distilled hLLM model doesn't generate the ranking. It scores the possible assignments, and a classical optimization algorithm constructs the ranking from those scores. The results are striking: 28 ms end-to-end inference and a reported 64× speedup, while maintaining ranking quality comparable to the teacher model. We've spent a lot of time figuring out how to make LLMs generate things more efficiently. But sometimes the bigger opportunity is to recognize that the problem we're solving has a structure that doesn't require generation at all. For ranking, that structure is a permutation. Once you recognize it, you can replace sequential decoding with a combinatorial optimization problem. Ultimately, the key is knowing when not to generate. #InformationRetrieval #Search #LLMRerankers Paper: https://lnkd.in/ebM_UY4G
250
8 Comments -
Amal Saad Alshehri
1K followers
Our recent paper, "Neural reranking for UK statutory retrieval: Provision-level evaluation and an open distilled model," has been published in the Artificial Intelligence and Law journal (Springer). In this study, we evaluate the effectiveness of neural reranking for retrieving UK statutory texts at the provision level. Alongside our findings, we have also released an open distilled model tailored specifically for legal information retrieval. The paper is Open Access and can be read in full here: https://lnkd.in/eEf7qXR8 #LegalTech #InformationRetrieval #UKLaw #OpenAccess
4
2 Comments -
Nishantha Ruwan
IWROBOTX Software Inc. • 2K followers
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values Outstanding Paper EMNLP 2025 The authors introduce ValueActionLens, a novel framework to assess how well large language models (LLMs) align their stated values with their value-informed actions. They construct a dataset of 14,784 value-informed actions covering 12 cultures and 11 social topics, rooted in the Schwartz theory of values. Experiments across multiple LLMs reveal substantial misalignment: models often declare a value (e.g., fairness) but choose an action that contradicts that value. These value-action gaps vary significantly by culture, topic, and model. The study underscores the risk of trusting stated values alone and calls for context-aware evaluations of value-action alignment in LLMs. https://lnkd.in/gJtgGUqK
-
Khoa Doan
VinUni-Illinois Smart Health… • 2K followers
Creativity with RL Trained Reasoning Model? Not quite yet. Despite strong performance (Pass@1) compared to the base LLM, RLVR reduces Pass@k at larger k, meaning the model solves fewer distinct problems given more attempts or the LLM loses diversity after RLVR training — the “reasoning boundary” shrinkage. This limits novel problem-solving and test time scaling. Why does this happen? And can we fix it? Introducing "negative interference", "winner-take-all", and "data curation" for RL Training of LLM in our newest work (highlights below): Phuc Minh Nguyen, Chinh La, Duy Ho Minh Nguyen, Nitesh Chawla, Binh T. Nguyen, Khoa D. Doan. The Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language Models. https://lnkd.in/gf7x4A_f -- our first joint work between the Center for Environmental Intelligence (CEI) at VinUniversity, and Prof. Nitesh Chawla at the University of Notre Dame and Lucy Family Institute for Data & Society (also our Honorary Professor VinUniversity). Highlights: - There exists **negative interference** during RLVL training, where a model update on problems in a batch lowers the log-likelihood of correct solutions for others out of batch, due to correlated gradients across different problems. This strength of neg. interference increases during training. - On-policy of RLVR likely reinforces problems highly solvable under the base LLM while neglecting problems with initially low likelihood of correct solutions. Neg. interference makes it worse: these examples capture most of the learning signals, resulting in winner-take-all learning behavior that narrows the solution diversity. Existing regularization e.g. clipping or reverse-KL cannot help mitigate this reasoning boundary shrinkage. - Solution? We propose SEFL, which selectively learns on problems with low likelihood of generating correct solutions, significantly improving the Pass@k performance of RLVR. Preprint: https://lnkd.in/gf7x4A_f Code: https://lnkd.in/gXg65Wbp
32
4 Comments -
Gayathri G
elsai • 4K followers
🔥 Hugging Face cooked again! They just dropped a free blog (more like a book) that dives deep into the no-BS reality of building SOTA models — covering the real decisions and trade-offs behind modern LLM research. I honestly haven’t seen any lab or researcher explain things this transparently. This is a gem. 💎 Syllabus: → Training compass: why → what → how → Every big model starts with a small ablation → Designing the model architecture → The art of data curation → The training marathon → Beyond base models — post-training in 2025 → Infrastructure — the unsung hero Skimming through it, this feels like their Ultrascale Playbook all over again — packed with practical insight instead of marketing fluff. I’ll definitely dig deeper and share some takeaways in the coming days. Link: https://lnkd.in/gxvnb_HX #HuggingFace #LLM #AIresearch #OpenSourceAI #MachineLearning
28
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content