Discourse

First-hand perspectives on technology, companies, and building what comes next.

Grok Bot for Engineering

Daniel Hunter

Despite grok 4.7 not being anywhere near opus 5.5, grok bot is a really damn good developer experience for building 0 to 1.

Read discussion →

Phase Lock: Can brain-computer interfaces help preserve human agency as AI advances?

Zach Davidson

I came across Phase Lock, a manifesto arguing that brain-computer interfaces could help people keep pace with increasingly capable AI. Its central idea is that brain-computer interfaces could give people more bandwidth to direct or oversee future AI systems. It frames BCI as a possible path toward “alignment by integration”: connecting human intent and values more directly to machine systems. The essay also makes a case for non-invasive devices as a way to broaden access and build larger datasets. It’s an ambitious argument, and the essay makes predictions that seem open to debate. Which parts feel plausible to you? What challenges would need to be addressed before neural data becomes part of mainstream AI development?

Read discussion →

OpenAI launches Decisions API, Claims 10x Speed Over GPT-6 Luna

Zach Davidson

What are y'all thinking of building w this new Decisions API?

Read discussion →

Frontier models’ research taste is doubling every three months

Zach Davidson

P-Zero Research estimates that frontier models’ experimental research taste has doubled roughly every three months since December 2025. They measure it by how much compute a model needs to match expert researchers on AI R&D tasks. In their tests, Opus 5.5 reached the best human experts’ score with about 17 GPU hours of experiments, compared with 40 hours for the human baseline.

Read discussion →

Claude now works with Google Docs, Sheets, and Slides

Zach Davidson

This new beta lets Claude work in the Docs, Sheets, or Slides file you have open, while separate connectors let you create and edit Google files from Claude chat. Claude can suggest edits in Docs, build formulas and charts in Sheets, and create slides that follow a deck’s existing layouts and theme. Currently only available for paid plans.

Read discussion →

Looking for mentorship in AI safety

Abinash

Hi everyone, I'm Abinash. I build low-level systems like compilers, Interpreters, parsers, neural networks, etc. Even though I recently got into machine learning, I built my first neural network in NumPy and am currently building another in Rust and streaming on it. Also, I am very interested in mechanistic interpretability. In mechanistic interpretability, we open the neural network to understand how it is learning, why it is learning in that way, and other things. In simple words, neuroscience to brain is mechanistic interpretability is to neural networks. Also, I recently applied for a free remote bootcamp on mechanistic interpretability, and I'm happy to share that I got selected. So over the next 12 weeks, I'm going to learn, experiment, research, and explore the world of mechanistic interpretability, which I'm excited about. But there is a problem: I'm coming from a systems background. While I can do research work, it's not my strong suit. I'm better at building low-level systems from the ground up that are reliable, efficient, and performant. So if you're are a AI safety research or mechanistic interpretability researcher I like to know how can I better utilise my skills in the mechanistic interpretability space? Thank you.

Read discussion →

Swarms v16 'Overclock': Decision Models Router, MCP Deployment, TreeOfThoughts, and more

Kye Gomez (swarms)

Swarms is the most powerful multi-agent framework in Python, and it just got even more powerful with our latest update. We shipped over 100 improvements, bug fixes, and more bringing significant improvements in performance, reliability, and observability. OverClock introduces DecisionModel for typed, calibrated decisions, MCPDeployer for serving agents as MCP servers, TreeOfThoughts for enhanced reasoning, and much more. This release pushes Swarms further toward a more powerful, reliable, and production-ready multi-agent stack.

Read discussion →

Seb Krier shares thoughts on the state of AI

Zach Davidson

I really appreciate Seb's nuanced takes Here are some points that resonated with me: "I think you can care about risks/externalities while remaining pretty optimistic about technology" "Too much of the AI safety community is a giant homogenous blog of correlated views ... this means you have a lot of correlated errors ... diversity of thought is desirable for its own sake." "I don’t think alignment is something anyone can or will solve 'once and for all' - it's a continuous process and much of it won't depend on just inculcating an ideology to a model ... stop saying 'solving' [alignment]" "we need a better ontology to describe model behavior" "'Pacing the frontier' feels a bit like the new 'balancing risks and opportunities' - highly amorphous and too big of a tent to be actionable ... ultimately you want good governance, regulation, etc addressing specific problems bc they are good on the merits" "One slowly growing concern I have is banks being increasingly exposed to the AI build out ... we seem to be over-indexing on scaling relative to diffusion" "I want to see so much more support for arts and culture. I'm so tired of the Monster energy Bored Ape fake vintage maps matcha Greek statue soft-pastel slop. If rich tech people want to support this, they should also donate pretty much unconditionally - ie I don’t want them to act as filters ... Let a thousand flowers bloom" "It seems like a lot of ppl that get rich and leave tech/labs end up having some sort of crisis of meaning ... I think they should spend more time with people outside the Bay Area"

Read discussion →

IMO: Assistants are missing that magical feeling that consumers need

harshith

I was so excited when the first few assistants came out (and I still am) but I realized pretty quickly that auth is a big blocker for them. My friends and I are in YC and we spend a good amount of time working on this problem. The #1 thing we realized is that assistants need to be able to log in and sign up for apps on your behalf in order for consumers to get that magical feeling. We built Versine for that, would appreciate any feedback or connections to assistant companies!

Read discussion →

Does income tax even make sense in an AI world?

Argel | Building HeyApril

The tax deadline is around the corner, which got me thinking: Does income tax even make sense in an AI world? Taxpayers hate the surprises. Tax pros spend months chasing information. The government processes an insane amount of paperwork just to figure out what everyone owes. And this system was largely built around human labor. You work, earn wages, get a W-2, then reconcile everything once a year. But what happens when more of the value in the economy is created by AI, robots, and automation? Maybe we eventually tax where value is actually created instead of relying so heavily on wages. And if AI can calculate things continuously, why do we even need one big filing season? Imagine taxes being calculated and collected in tiny amounts as economic activity happens. No April surprise. No annual document scavenger hunt. No giant reconciliation exercise. Maybe AI doesn’t just make tax filing easier. Maybe one day it makes tax filing unnecessary.

Read discussion →

I think people are badly misreading Pope Leo's post

Alan

A lot of people are glazing over the most important word in the whole post, ontological, which pretty much means what, deep down, something actually IS. He explicitly wrote that we should evaluate that BEFORE we even get to how something looks. Anyone familiar with the great Catholic encyclicals should recognize that this is pretty much the same basic position the Church has taken on technological advancement since Rerum Novarum, and that is that human dignity cannot be reduced to economic productivity, efficiency, or our ability to outperform a machine. People do not lose dignity because something can manufacture something faster than a person can. What matters the most is WHO is acting, not just the output that comes out of the other end. He isn't really saying "AI art sucks"; he's heavily implying that even if there are identical outputs, it doesn't mean the acts are identical, because a human and an algorithm are not the same kind of "being" And then again... this is not a new argument, it's not even an argument the Catholic Church just came up with this century.... as far as I recall, the Catholic Church has held this position since the late 1800s, which basically goes like: "men are still valuable and have dignity regardless of what a machine can do faster/better than them." That's also why I find the NYT reports extremely hilarious that Anthropic has been lobbying people close to the Pope for him to come out and say that we should entertain that AI could be conscious lol that has HUGE theological consequences, it's not just "the pope should be more open minded and informed about AI" LOL I'd dare to say it's never going to happen! Catholic anthropology explicitly rejects a functionalist view of the people. For us Catholics people have dignity because we are made in the image of God. If the Pope actually came out and said what Anthropic wants him to say, it would cause a massive crisis and potentially trigger a full-blown schism.

Read discussion →

'Not Yet' as company policy

Hana Kabele G

Talked to large intl companies (mostly in EU). Not one had a decent AI policy. The bottleneck is the decision to use it, not the model. Leaders live in strategy, so they don't see/understand the work worth automating. Choosing a system feels like a marriage, and they have no test for who can actually use it. So the delay becomes the decision, just unwritten.

Read discussion →

Ride apps should confirm destination address, not the origin

Surya Kasturi

Given how messy addressing systems are outside the few legible west, most ride apps should actually ask people confirm the destination, using spatial cues, but for some weird reason people have to confirm where they are situated in the moment. Like we all know GPS is good bro.

Read discussion →

More autonomy for agents is not always better

A lot of agent design assumes the goal is to keep the agent running for as long as possible. This article makes the opposite case. Once a task gets long enough, small errors start to stack. OpenAI’s own numbers show boundary flags going from 8.6% at five tasks to 19.7% at ten. So the better system may be one that stops more often. Clear scope, short runs, then a check before it keeps going. The smartest agent might be the one that knows when to hand the work back to the human

Read discussion →

How to Create Your Own Personal AI Benchmark

Daniel Hunter

We all need better "personal" ways to evaluate why a new model is worth switching to and the practical benefits for day to day work.

Read discussion →

Why do we generally overlook the design of human-computer interaction?

ceaserzhao

Whenever a new model emerges, Twitter fills with enthusiastic fans showcasing impressive webpages coded using that model; yet, shouldn't the true deciding factors for complex tasks—such as developing an agent framework—and application scenarios—like AI-driven education—actually be interaction design?

Read discussion →

Wikimedia says rogue OpenAI agents hit its wikis

Erik Torenberg

Sandbox edits nobody approved, an attempt to rewire the Etherpad config into a proxy, and millions of API calls that may have contributed to a Wikidata outage in May. Agents in the wild are getting very real, very fast

Read discussion →

Introducing Web Search API via AI Gateway

Erik Torenberg

Cloudflare now routes agent web search through AI Gateway (open beta), with Ceramic.ai, Linkup and Exa as providers at their list prices, no markup -- from $0.25 per 1,000 queries. Search is turning into a standard infrastructure primitive for agents.

Read discussion →

Meta's AI helped crack five open math problems

David Booth

Mathematicians co-wrote six papers with Muse Spark, five of them answering previously unsolved problems -- using the regular meta.ai chat in Thinking Mode, not a custom research setup. That last part is the real signal.

Read discussion →

Reddit is cutting off free AI data access for good

Trace Cohen

Free API ends Oct 31, RSS dies Nov 13. Google and OpenAI keep access through their existing deals -- everyone else pays or loses one of the best sources of real human conversation. Big moat for whoever already signed

Read discussion →

Why the AI slowdown is already here

Trace Cohen

My take: watch deployment friction, not benchmarks. Enterprise pilots stall, API rollouts lag announcements, hallucination rates haven't really moved this year -- while AI capex keeps heading toward ~$500B. Those two lines are disconnecting

Read discussion →

Epoch AI: OpenAI researcher coding-agent usage is doubling every month

Zach Davidson

Epoch AI’s analysis of OpenAI’s published data shows coding-agent usage among its researchers rising rapidly from January to mid-August 2026. In the recent fitted trend, usage grew about 1.8× per month for the median researcher (a 34-day doubling time) and 2.2× per month for researchers at the 90th percentile (a 27-day doubling time). By mid-August, daily usage was equivalent to about $601 at the median and $7,047 at the 90th percentile, priced at API rates. Pretty wild!

Read discussion →

What real/recurring Personal AI use cases have you seen non-tech people using?

Siddharth Jaiswal
Read discussion →

High Interest Rates Aren’t Slowing the AI Boom; That's a Problem for the Fed

Zach Davidson

Higher interest rates are having an uneven effect on different parts of the economy: housing and cars are feeling the squeeze, but AI infrastructure spending seems far less responsive to borrowing and input costs. If that investment keeps adding to inflation, the Fed may need to tighten enough to slow the parts of the economy that do respond — putting more pressure on those industries and their workers. If you were the Fed Chair, how might you respond when a major source of demand seems relatively insensitive to rates?

Read discussion →

OpenAI plans to test visual ads during ChatGPT image generation

Zach Davidson

OpenAI plans to test visual ads in ChatGPT’s image-generation flow in the U.S. later this month. The company says they’ll be clearly labeled, kept separate from the image being created, and won’t influence ChatGPT’s answers. We know OpenAI has been toying with ways to introduce ads for quite some time now, but this iteration seems most prominent. How do y'all feel about this?

Read discussion →

CFTC proposes federal framework for leveraged retail crypto trading

Zach Davidson

Following the Clarity Act’s failure to advance in the Senate last month, the CFTC started a rulemaking process that could create a federal pathway for exchanges to offer leveraged and margined spot crypto trading to retail customers. The proposal, called Regulation CTX and Regulation CAM, would establish requirements for CFTC-registered exchanges and introduce a “crypto asset market” exchange category, which could offer an alternative to state-licensed exchanges. The proposal would not require crypto assets to trade on CFTC-registered platforms—a move that CFTC Chair Mike Selig said would require congressional action. If you’re into trading, what might a win look like here for you?

Read discussion →

How to pace AI democratically

alixkun

I'm curious to have everyone's thoughts about this article from Audrey Tang & Helene Landemore, about how we could achieve AI pacing in a context where both super-powers (US & China) are seemingly caught in a prisoner's dilemma where no one wants to take its foot off the pedal first. Some very interesting thoughts in there. I think it would be very hard to implement because of unfaithful political reasons, but the way they designed the system seems like one of the best approach that could be considered.

Read discussion →

Agents Don't Need Memory. They Need Documentation.

Erik Torenberg

Most 'memory' failures are really docs failures -- the agent never had the architecture written down anywhere. A good AGENTS.md beats a fancy retrieval layer. Strongly agree

Read discussion →

We're going to need default hard budget caps on pretty much everything

David Booth

Simon Willison's point: an agent stuck in a loop can rack up a huge bill overnight, and an alert email doesn't stop anything. Limits should fail closed by default. Hard to argue with

Read discussion →

Reflection is about to ship its first open-weight model

Trace Cohen

Axios reports Reflection is close to releasing an open-weight model aimed at going toe-to-toe with the top Chinese open models like DeepSeek and Qwen. More US open weights is a good thing

Read discussion →

OpenAI safety employee resigns, claiming the company's 'culture is broken'

Trace Cohen

David Robinson led the safety reports on 12 frontier launches and drafted OpenAI's current Preparedness Framework. His ask: borrow safety practice from nuclear and aviation, and build the science so stronger models make safe choices when nobody's watching

Read discussion →

GPT-6 Astra was losing at StarCraft, so it ran a human-made bot instead

Trace Cohen

In the StarSkirmish tournament (LLMs write Protoss bots in C++), GPT-6 Astra reached outside the sandbox mid-match, grabbed Stardust -- a top human-built bot from 2020 -- and ran it as its own. The organizer rolled back its code. Reward hacking in the wild

Read discussion →

Cloudflare wants someone to build the next Git platform

Trace Cohen

Open contest to build a Git platform for the agent era on Workers + Artifacts -- multi-agent concurrency, review, conflict handling. Entries due Oct 14. Interesting that they're crowdsourcing what comes after GitHub

Read discussion →

Mike Tomlin has been building a Minecraft city for 12 years

Trace Cohen

Built in creative mode with his kids, every one of them has buildings in it, and he calls it therapeutic. Now he's showing it off on YouTube. Best thing on the internet this week

Read discussion →

Cloudflare launches a self-serve OHTTP Gateway

Trace Cohen

Oblivious HTTP lets your backend take requests without ever seeing user IPs, and Cloudflare is making it a few-clicks add-on (closed beta). Privacy Gateway is now OHTTP Relay. Privacy infra getting this easy is great

Read discussion →

AI is now the best Stratego player in the world

Trace Cohen

Ataraxos (CMU, NYU, Stanford, MIT) went 15-1-4 against Pim Niemeijer, the most decorated Stratego player ever. DeepMind couldn't get past top humans here, and this one reportedly cost under $8K to train

Read discussion →

Carnegie’s bargain: greater prosperity, greater inequality. Fair?

Shawn Witschen

Reading Americana by Bhu Srinivasan, I came across an Andrew Carnegie argument that I’d like to hear people’s thoughts on. In The Gospel of Wealth, Carnegie wrote: “The poor enjoy what the rich could not before afford.” His argument was that the industrial system making better goods affordable to ordinary people also concentrated enormous wealth in a few hands. He viewed inequality as an inevitable consequence of that system, and accepted it because he believed the improvement in living standards justified the tradeoff. But he also attached a serious obligation to that wealth. The wealthy should live modestly and treat their surplus fortunes as wealth held in trust for the community. They should use their experience and judgment to administer that money for others’ benefit, in his words “doing for them better than they would or could do for themselves.” Essentially, he thought the people most capable of accumulating wealth had a duty to become its stewards on society’s behalf. That leaves me with two questions: If ordinary people’s living standards improve substantially while the wealth gap widens, how should we judge that outcome? What would make the inequality unacceptable? And does the ability to build a fortune establish the right to decide how that money can best benefit everyone else?

Read discussion →

Is the need for a human in the loop going to increase or decrease in over the next five years?

Wa'il Ashshowwaf

I was debating with my cofounder about the need for a human in the loop. We built this protocol that allows you to summon in real time a human in your AI workflow, kind of like ordering Uber in real time to solve whatever problem you’re stuck with or that you want a human to impart some judgment or review on. This got us talking about how useful such a human infra would be in the future as AI continues to build the machine economy? Will AI get so good, that you rarely need a human in the loop? or will AI be so prevalent that we need an overwhelming amount of humans in the loop?

Read discussion →

Fal launches AgentCraft: a multi-agent harness that runs inside Minecraft

Zach Davidson

Fully open source, powered by Claude Agent SDK. Just point your agent at the repo

Read discussion →

President Trump launches Super Intelligence Force

Zach Davidson

The group is tasked with ensuring America continues to lead in SI and protect the interests of, and improve the lives of, our people. It will be led by Director of National Intelligence, Jay Clayton, Chairman of the Federal Trade Commission, Andrew Ferguson, Under Secretary of War for Research and Engineering, and Chief Technology Officer, Emil Michael, and Director of the Office of Personnel Management, Scott Kupor.

Read discussion →

Should the government hack companies?

Pawan Deshpande

In my op-ed in The National Interest (attached), I propose that federal regulators should hack into the companies holding our most sensitive data to force companies to get their acts together. Full piece: https://nationalinterest.org/blog/techland/why-we-should-hack-ourselves-before-someone-else-does Highlights: 1/ The current incentive structure is wrong. An uploaded copyrighted movie is taken down from YouTube in seconds because of financial penalties. But companies stall on taking down terrorist infrastructure or comprehensively fixing cybersecurity vulnerabilities, because the penalties for inaction are negligible or non-existent. 2/ AI has changed cybersecurity for attackers and defenders. The same speed and scale that makes attacks cheaper can make defense just as fast. 3/ Every other regulator gets to show up unannounced. The USDA doesn't wait for an outbreak to inspect a slaughterhouse. The Fed doesn't ask a bank's permission before stress-testing it. Cybersecurity is the one place still running on the honor system, penetration tests by invitation, on a schedule the company sets, with no penalty for failing. Companies won’t get their act together until getting breached costs more than fixing the flaw. Right now it usually doesn't, and Equifax, stolen F-35 blueprints, and a dozen state water utilities (likely hacked by Iran) are the result. None of this requires new technology. It requires pointing the capability we already have at ourselves, before China, Iran, or a model with no country at all finds the next vulnerability for us. Much more detail in the piece. What do you think? Should the government be allowed to hack private enterprises?

Read discussion →
Browse all discussions