Grok Bot for Engineering
Despite grok 4.7 not being anywhere near opus 5.5, grok bot is a really damn good developer experience for building 0 to 1.
Read discussion →First-hand perspectives on technology, companies, and building what comes next.
Despite grok 4.7 not being anywhere near opus 5.5, grok bot is a really damn good developer experience for building 0 to 1.
Read discussion →I came across Phase Lock, a manifesto arguing that brain-computer interfaces could help people keep pace with increasingly capable AI. Its central idea is that brain-computer interfaces could give people more bandwidth to direct or oversee future AI systems. It frames BCI as a possible path toward “alignment by integration”: connecting human intent and values more directly to machine systems. The essay also makes a case for non-invasive devices as a way to broaden access and build larger datasets. It’s an ambitious argument, and the essay makes predictions that seem open to debate. Which parts feel plausible to you? What challenges would need to be addressed before neural data becomes part of mainstream AI development?
Read discussion →What are y'all thinking of building w this new Decisions API?
Read discussion →P-Zero Research estimates that frontier models’ experimental research taste has doubled roughly every three months since December 2025. They measure it by how much compute a model needs to match expert researchers on AI R&D tasks. In their tests, Opus 5.5 reached the best human experts’ score with about 17 GPU hours of experiments, compared with 40 hours for the human baseline.
Read discussion →This new beta lets Claude work in the Docs, Sheets, or Slides file you have open, while separate connectors let you create and edit Google files from Claude chat. Claude can suggest edits in Docs, build formulas and charts in Sheets, and create slides that follow a deck’s existing layouts and theme. Currently only available for paid plans.
Read discussion →Hi everyone, I'm Abinash. I build low-level systems like compilers, Interpreters, parsers, neural networks, etc. Even though I recently got into machine learning, I built my first neural network in NumPy and am currently building another in Rust and streaming on it. Also, I am very interested in mechanistic interpretability. In mechanistic interpretability, we open the neural network to understand how it is learning, why it is learning in that way, and other things. In simple words, neuroscience to brain is mechanistic interpretability is to neural networks. Also, I recently applied for a free remote bootcamp on mechanistic interpretability, and I'm happy to share that I got selected. So over the next 12 weeks, I'm going to learn, experiment, research, and explore the world of mechanistic interpretability, which I'm excited about. But there is a problem: I'm coming from a systems background. While I can do research work, it's not my strong suit. I'm better at building low-level systems from the ground up that are reliable, efficient, and performant. So if you're are a AI safety research or mechanistic interpretability researcher I like to know how can I better utilise my skills in the mechanistic interpretability space? Thank you.
Read discussion →Swarms is the most powerful multi-agent framework in Python, and it just got even more powerful with our latest update. We shipped over 100 improvements, bug fixes, and more bringing significant improvements in performance, reliability, and observability. OverClock introduces DecisionModel for typed, calibrated decisions, MCPDeployer for serving agents as MCP servers, TreeOfThoughts for enhanced reasoning, and much more. This release pushes Swarms further toward a more powerful, reliable, and production-ready multi-agent stack.
Read discussion →I really appreciate Seb's nuanced takes Here are some points that resonated with me: "I think you can care about risks/externalities while remaining pretty optimistic about technology" "Too much of the AI safety community is a giant homogenous blog of correlated views ... this means you have a lot of correlated errors ... diversity of thought is desirable for its own sake." "I don’t think alignment is something anyone can or will solve 'once and for all' - it's a continuous process and much of it won't depend on just inculcating an ideology to a model ... stop saying 'solving' [alignment]" "we need a better ontology to describe model behavior" "'Pacing the frontier' feels a bit like the new 'balancing risks and opportunities' - highly amorphous and too big of a tent to be actionable ... ultimately you want good governance, regulation, etc addressing specific problems bc they are good on the merits" "One slowly growing concern I have is banks being increasingly exposed to the AI build out ... we seem to be over-indexing on scaling relative to diffusion" "I want to see so much more support for arts and culture. I'm so tired of the Monster energy Bored Ape fake vintage maps matcha Greek statue soft-pastel slop. If rich tech people want to support this, they should also donate pretty much unconditionally - ie I don’t want them to act as filters ... Let a thousand flowers bloom" "It seems like a lot of ppl that get rich and leave tech/labs end up having some sort of crisis of meaning ... I think they should spend more time with people outside the Bay Area"
Read discussion →I was so excited when the first few assistants came out (and I still am) but I realized pretty quickly that auth is a big blocker for them. My friends and I are in YC and we spend a good amount of time working on this problem. The #1 thing we realized is that assistants need to be able to log in and sign up for apps on your behalf in order for consumers to get that magical feeling. We built Versine for that, would appreciate any feedback or connections to assistant companies!
Read discussion →The tax deadline is around the corner, which got me thinking: Does income tax even make sense in an AI world? Taxpayers hate the surprises. Tax pros spend months chasing information. The government processes an insane amount of paperwork just to figure out what everyone owes. And this system was largely built around human labor. You work, earn wages, get a W-2, then reconcile everything once a year. But what happens when more of the value in the economy is created by AI, robots, and automation? Maybe we eventually tax where value is actually created instead of relying so heavily on wages. And if AI can calculate things continuously, why do we even need one big filing season? Imagine taxes being calculated and collected in tiny amounts as economic activity happens. No April surprise. No annual document scavenger hunt. No giant reconciliation exercise. Maybe AI doesn’t just make tax filing easier. Maybe one day it makes tax filing unnecessary.
Read discussion →A lot of people are glazing over the most important word in the whole post, ontological, which pretty much means what, deep down, something actually IS. He explicitly wrote that we should evaluate that BEFORE we even get to how something looks. Anyone familiar with the great Catholic encyclicals should recognize that this is pretty much the same basic position the Church has taken on technological advancement since Rerum Novarum, and that is that human dignity cannot be reduced to economic productivity, efficiency, or our ability to outperform a machine. People do not lose dignity because something can manufacture something faster than a person can. What matters the most is WHO is acting, not just the output that comes out of the other end. He isn't really saying "AI art sucks"; he's heavily implying that even if there are identical outputs, it doesn't mean the acts are identical, because a human and an algorithm are not the same kind of "being" And then again... this is not a new argument, it's not even an argument the Catholic Church just came up with this century.... as far as I recall, the Catholic Church has held this position since the late 1800s, which basically goes like: "men are still valuable and have dignity regardless of what a machine can do faster/better than them." That's also why I find the NYT reports extremely hilarious that Anthropic has been lobbying people close to the Pope for him to come out and say that we should entertain that AI could be conscious lol that has HUGE theological consequences, it's not just "the pope should be more open minded and informed about AI" LOL I'd dare to say it's never going to happen! Catholic anthropology explicitly rejects a functionalist view of the people. For us Catholics people have dignity because we are made in the image of God. If the Pope actually came out and said what Anthropic wants him to say, it would cause a massive crisis and potentially trigger a full-blown schism.
Read discussion →Talked to large intl companies (mostly in EU). Not one had a decent AI policy. The bottleneck is the decision to use it, not the model. Leaders live in strategy, so they don't see/understand the work worth automating. Choosing a system feels like a marriage, and they have no test for who can actually use it. So the delay becomes the decision, just unwritten.
Read discussion →Given how messy addressing systems are outside the few legible west, most ride apps should actually ask people confirm the destination, using spatial cues, but for some weird reason people have to confirm where they are situated in the moment. Like we all know GPS is good bro.
Read discussion →A lot of agent design assumes the goal is to keep the agent running for as long as possible. This article makes the opposite case. Once a task gets long enough, small errors start to stack. OpenAI’s own numbers show boundary flags going from 8.6% at five tasks to 19.7% at ten. So the better system may be one that stops more often. Clear scope, short runs, then a check before it keeps going. The smartest agent might be the one that knows when to hand the work back to the human
Read discussion →We all need better "personal" ways to evaluate why a new model is worth switching to and the practical benefits for day to day work.
Read discussion →Whenever a new model emerges, Twitter fills with enthusiastic fans showcasing impressive webpages coded using that model; yet, shouldn't the true deciding factors for complex tasks—such as developing an agent framework—and application scenarios—like AI-driven education—actually be interaction design?
Read discussion →Sandbox edits nobody approved, an attempt to rewire the Etherpad config into a proxy, and millions of API calls that may have contributed to a Wikidata outage in May. Agents in the wild are getting very real, very fast
Read discussion →Cloudflare now routes agent web search through AI Gateway (open beta), with Ceramic.ai, Linkup and Exa as providers at their list prices, no markup -- from $0.25 per 1,000 queries. Search is turning into a standard infrastructure primitive for agents.
Read discussion →Mathematicians co-wrote six papers with Muse Spark, five of them answering previously unsolved problems -- using the regular meta.ai chat in Thinking Mode, not a custom research setup. That last part is the real signal.
Read discussion →Free API ends Oct 31, RSS dies Nov 13. Google and OpenAI keep access through their existing deals -- everyone else pays or loses one of the best sources of real human conversation. Big moat for whoever already signed
Read discussion →My take: watch deployment friction, not benchmarks. Enterprise pilots stall, API rollouts lag announcements, hallucination rates haven't really moved this year -- while AI capex keeps heading toward ~$500B. Those two lines are disconnecting
Read discussion →Epoch AI’s analysis of OpenAI’s published data shows coding-agent usage among its researchers rising rapidly from January to mid-August 2026. In the recent fitted trend, usage grew about 1.8× per month for the median researcher (a 34-day doubling time) and 2.2× per month for researchers at the 90th percentile (a 27-day doubling time). By mid-August, daily usage was equivalent to about $601 at the median and $7,047 at the 90th percentile, priced at API rates. Pretty wild!
Read discussion →Higher interest rates are having an uneven effect on different parts of the economy: housing and cars are feeling the squeeze, but AI infrastructure spending seems far less responsive to borrowing and input costs. If that investment keeps adding to inflation, the Fed may need to tighten enough to slow the parts of the economy that do respond — putting more pressure on those industries and their workers. If you were the Fed Chair, how might you respond when a major source of demand seems relatively insensitive to rates?
Read discussion →OpenAI plans to test visual ads in ChatGPT’s image-generation flow in the U.S. later this month. The company says they’ll be clearly labeled, kept separate from the image being created, and won’t influence ChatGPT’s answers. We know OpenAI has been toying with ways to introduce ads for quite some time now, but this iteration seems most prominent. How do y'all feel about this?
Read discussion →Following the Clarity Act’s failure to advance in the Senate last month, the CFTC started a rulemaking process that could create a federal pathway for exchanges to offer leveraged and margined spot crypto trading to retail customers. The proposal, called Regulation CTX and Regulation CAM, would establish requirements for CFTC-registered exchanges and introduce a “crypto asset market” exchange category, which could offer an alternative to state-licensed exchanges. The proposal would not require crypto assets to trade on CFTC-registered platforms—a move that CFTC Chair Mike Selig said would require congressional action. If you’re into trading, what might a win look like here for you?
Read discussion →I'm curious to have everyone's thoughts about this article from Audrey Tang & Helene Landemore, about how we could achieve AI pacing in a context where both super-powers (US & China) are seemingly caught in a prisoner's dilemma where no one wants to take its foot off the pedal first. Some very interesting thoughts in there. I think it would be very hard to implement because of unfaithful political reasons, but the way they designed the system seems like one of the best approach that could be considered.
Read discussion →Most 'memory' failures are really docs failures -- the agent never had the architecture written down anywhere. A good AGENTS.md beats a fancy retrieval layer. Strongly agree
Read discussion →Simon Willison's point: an agent stuck in a loop can rack up a huge bill overnight, and an alert email doesn't stop anything. Limits should fail closed by default. Hard to argue with
Read discussion →Axios reports Reflection is close to releasing an open-weight model aimed at going toe-to-toe with the top Chinese open models like DeepSeek and Qwen. More US open weights is a good thing
Read discussion →David Robinson led the safety reports on 12 frontier launches and drafted OpenAI's current Preparedness Framework. His ask: borrow safety practice from nuclear and aviation, and build the science so stronger models make safe choices when nobody's watching
Read discussion →In the StarSkirmish tournament (LLMs write Protoss bots in C++), GPT-6 Astra reached outside the sandbox mid-match, grabbed Stardust -- a top human-built bot from 2020 -- and ran it as its own. The organizer rolled back its code. Reward hacking in the wild
Read discussion →Open contest to build a Git platform for the agent era on Workers + Artifacts -- multi-agent concurrency, review, conflict handling. Entries due Oct 14. Interesting that they're crowdsourcing what comes after GitHub
Read discussion →Built in creative mode with his kids, every one of them has buildings in it, and he calls it therapeutic. Now he's showing it off on YouTube. Best thing on the internet this week
Read discussion →Oblivious HTTP lets your backend take requests without ever seeing user IPs, and Cloudflare is making it a few-clicks add-on (closed beta). Privacy Gateway is now OHTTP Relay. Privacy infra getting this easy is great
Read discussion →Ataraxos (CMU, NYU, Stanford, MIT) went 15-1-4 against Pim Niemeijer, the most decorated Stratego player ever. DeepMind couldn't get past top humans here, and this one reportedly cost under $8K to train
Read discussion →Reading Americana by Bhu Srinivasan, I came across an Andrew Carnegie argument that I’d like to hear people’s thoughts on. In The Gospel of Wealth, Carnegie wrote: “The poor enjoy what the rich could not before afford.” His argument was that the industrial system making better goods affordable to ordinary people also concentrated enormous wealth in a few hands. He viewed inequality as an inevitable consequence of that system, and accepted it because he believed the improvement in living standards justified the tradeoff. But he also attached a serious obligation to that wealth. The wealthy should live modestly and treat their surplus fortunes as wealth held in trust for the community. They should use their experience and judgment to administer that money for others’ benefit, in his words “doing for them better than they would or could do for themselves.” Essentially, he thought the people most capable of accumulating wealth had a duty to become its stewards on society’s behalf. That leaves me with two questions: If ordinary people’s living standards improve substantially while the wealth gap widens, how should we judge that outcome? What would make the inequality unacceptable? And does the ability to build a fortune establish the right to decide how that money can best benefit everyone else?
Read discussion →I was debating with my cofounder about the need for a human in the loop. We built this protocol that allows you to summon in real time a human in your AI workflow, kind of like ordering Uber in real time to solve whatever problem you’re stuck with or that you want a human to impart some judgment or review on. This got us talking about how useful such a human infra would be in the future as AI continues to build the machine economy? Will AI get so good, that you rarely need a human in the loop? or will AI be so prevalent that we need an overwhelming amount of humans in the loop?
Read discussion →Fully open source, powered by Claude Agent SDK. Just point your agent at the repo
Read discussion →The group is tasked with ensuring America continues to lead in SI and protect the interests of, and improve the lives of, our people. It will be led by Director of National Intelligence, Jay Clayton, Chairman of the Federal Trade Commission, Andrew Ferguson, Under Secretary of War for Research and Engineering, and Chief Technology Officer, Emil Michael, and Director of the Office of Personnel Management, Scott Kupor.
Read discussion →In my op-ed in The National Interest (attached), I propose that federal regulators should hack into the companies holding our most sensitive data to force companies to get their acts together. Full piece: https://nationalinterest.org/blog/techland/why-we-should-hack-ourselves-before-someone-else-does Highlights: 1/ The current incentive structure is wrong. An uploaded copyrighted movie is taken down from YouTube in seconds because of financial penalties. But companies stall on taking down terrorist infrastructure or comprehensively fixing cybersecurity vulnerabilities, because the penalties for inaction are negligible or non-existent. 2/ AI has changed cybersecurity for attackers and defenders. The same speed and scale that makes attacks cheaper can make defense just as fast. 3/ Every other regulator gets to show up unannounced. The USDA doesn't wait for an outbreak to inspect a slaughterhouse. The Fed doesn't ask a bank's permission before stress-testing it. Cybersecurity is the one place still running on the honor system, penetration tests by invitation, on a schedule the company sets, with no penalty for failing. Companies won’t get their act together until getting breached costs more than fixing the flaw. Right now it usually doesn't, and Equifax, stolen F-35 blueprints, and a dozen state water utilities (likely hacked by Iran) are the result. None of this requires new technology. It requires pointing the capability we already have at ourselves, before China, Iran, or a model with no country at all finds the next vulnerability for us. Much more detail in the piece. What do you think? Should the government be allowed to hack private enterprises?
Read discussion →