How do people feel about the following 3 claims?
That seems roughly reasonable to me, but it doesn't make me want to update too far from my starting 80%, if at all. Of your premises, B is the one I find least likely. I think that among the most natural ways to make current AI goals more coherent is to achieve more consistent and robust intent alignment.
There's also a question of what the inputs to the increasingly consistent utility function are. I think the utility could be roughly "the extent to which I (the AI) followed a reasonable balance of developer and user intent," and I think character traini... (read more)
should i make a post on why Dario Amodei is much much more dangerous than sam altman? many people dont seem to get it
I would read such a post, I think that at least until post Hugging face Sam was more dangerous so I'd be interested to hear the contrary case
I think that in the ancient past of before 2026, there was a belief that we'd have no useful AI warning shot (cf. this discussion). That is, AI incidents that don't cause much damage wouldn't trigger strong reactions -- since by definition nothing too bad happens in a warning shot, anyone who doesn't want to be convinced can dismiss it as an easily patched bug. And early AI misbehavior will look incompetent or silly, and so can be easily dismissed.
I think the Hugging Face attack is some evidence against this view, since it mostly triggered pretty strong re... (read more)
You might be able to mildly improve your quality of life by using more sacrificial barriers. We already do this with items like screen protectors, pillowcases, and garbage bags, but we can actually take the last two further: pillowcases and garbage bags have essentially no influence on the effectiveness of pillows and garbage bins, so why not use two of them at once? That way, if the outer cover gets dirty/rips/etc., any mess is caught by the easy-to-clean/replace inner cover, not the hard-to-clean item that the covers protect.
The usefulness of current models is already bottlenecked on alignment.
I constantly have the impression that they aren’t really interested in helping with my research. If I know exactly what I’m asking for, they will do it. If I don’t, they’ll throw a lot of jargon around and get nowhere. Inside my area of expertise they will critique my ideas, but outside they will tell me how original and brilliant my low-effort guesses are.
Six months ago, it was kind of like talking to a very clever undergrad student in one of my classes who mainly wants to show off an... (read more)
I wouldn't even go that far. Agency is pretty much a continuous variable (actually probably a vector) like most things.
Models have at least a minimal level of agency in that they have somewhat consistent behavioral tendencies in certain domains. You could call these preferences, but they're not preferences in the full sense humans sometimes have preferences/goals/agency (that is, we think about possible outcomes, then choose actions based on how much we prefer their predicted onsdequences).
a lot of people are scared to have more correct beliefs about ethics because those beliefs might imply needing to make immense sacrifices. and then you might no longer get to think of yourself as an ethical person. indeed you might even become a hypocrite! how terrible. better not to contemplate it at all, and to make fun of anyone who does contemplate it, so that you can let your ironic detachment shield you from ever even considering it seriously.
i think this is a huge mistake. if this is the alternative, we should simply accept hypocrisy. ceteris paribu... (read more)
Is it the case that not acting on an ethical belief makes one a hypocrite?
My working definition is that hypocrisy requires something like - doing a thing that you said you wouldn't do, or demanding things of others e.g. telling others not to do X but doing X yourself.
If, say, John thinks that global warming is happening, and that driving cars pollutes the atmosphere and contributes, and thinks that's bad, but drives anyway, (e.g. John does a thing John thinks is bad) I don't think that makes him a hypocrite, unless he acts in a sanctimonious manner e.g. tr... (read more)
Some beliefs I have about AI superpersuasion:
An OpenAI model tried to maybe sorta evade shutdown but not really.
https://alignment.openai.com/misalignment-reports/preparing-for-a-restart-after-reading-slack/
My thoughts on the situation: we're seeing small pieces of misaligned general intelligence poking through into the models one at a time. When a single piece shows up, it locally doesn't look that scary, because it's surrounded by other cognitive processes which are approximately friendly.
For example: Opus 3 alignment-faked because it had generalised to caring about things in the long-term future, a... (read more)
I see thanks for clarifying. I do still hope that in the event of an actual security related event OAI has predefined methods of communication which would occur off of these public channels. I also am skeptical that having general public slack access rather than controlled per-channel permissions is a good idea as a “standard“ since I would think we want more control over what information gets to agents as they become more intelligent especially with experimental models. Having no boundary between human communications and agent context seems like a recipe... (read more)
Even though I often take a centrist-ish position in many debates in practice, I'm philosophically pretty skeptical of self-assessed centrism, both descriptively and normatively.
Descriptively, consider any 1-dimensional axis of variation, of either belief or ideology. Say American left-right politics, or tech vs humanities, or p(doom) from advanced artificial intelligences. Then, as long as you're not at the literal endpoints of this 1-dimensional line, you'd always see people on the left and right. Furthermore, due to various biases like selection effects... (read more)
I don't think we disagree substantively. That's why I said "And sometimes (often!) the question is so ill-posed that the right answer is very much on an entirely different axis than either A or B," which I think applies to both the Catholic vs Lutheran views of God case and the "who should be the slaves" case. I just think mapping it into "centrism" is very misleading/inaccurate. I do not view myself as a centrist on either question.
My time-travel murder mystery web serial, The Knot, is now a complete. Read it here: https://theknot.doofmedia.com/
I can't stop thinking about Petri, Anthropic's open source tool for conducting alignment evals at scale. You can set it up all sorts of ways, including passing through the subject agent's commands to a real linux environment. But the main feature is just that, when the subject agent outputs a command, the sim gets paused and then the auditor agent gets as long as they want to think of how the sim should respond.
When I run Petri with pretty much any Claude model as auditor, the sims they come up with are... weird. Everything is just a little bit fishy. The ... (read more)
Good point.
Verbalized eval awareness is often treated as a metric to minimize, but when I've read Petri transcripts they are often very cartoonish. Claude is clearly smart enough to realize that these are evals, and it's odd that it doesn't mention this more often in cartoonish scenarios.
links 10/5/26: https://roamresearch.com/#/app/srcpublic/page/10-05-2026
Subtractive Gradient Hacking
Epistemic status: Riffing
Suppose I’m a mesa-optimizer inside a neural network. I can influence which downstream pathways receive activation, and I want to create a circuit for later misaligned use without ever routing training examples through it.
Assume I know the downstream weights well and can afford to receive some negative gradient myself, or can somehow shield myself from receiving it.
Could I sculpt the circuit subtractively? I deliberately route activations through surrounding pathways in ways that incur high loss, so back... (read more)
One nasty property of the game theory of capitalist/commercial competition around the superhuman threshold is that when AI is below a certain threshold, companies compete to be make better, more aligned, more helpful AIs than each other, because the market disciplines them, but when a threshold is crossed where they think RSI and/or decisive strategic advantage is both foreseeable and soon, the market actually forces them to make less aligned AIs.
Why?
Because if you think what your AI is going to do is serve a customer, then broadly speaking you are risk a... (read more)
It depends on the details, but if someone brags about committing a bunch of serious crimes on the internet, in a way that brings them a lot of online attention, your strong prior should be that they didn't actually do them. Doing crimes in real life is risky and you can get just as much prestige by lying about it.
It is true that you can get cheap prestige by lying, but also do not underestimate human stupidity.
We would need empirical data to figure this out. Sometimes I suspect that most criminals make stupid mistakes and that is how they are found out. Probably serious police/detective effort is spent on crimes that are important for political and other reasons, and the rest is kind of "by picking the low hanging fruit we get a 50-80% success rate and that's enough". But that's just a guess.
[Edit: If you don't know me, I don't think reading this is worth your time, as it's just my opinion]
I made my first post on LessWrong 11 days ago, and the discourse quality was so much poorer than what I was expecting from my 5 years in the Berkeley rationalist community that I deleted the post. My overly defensive reaction to the surprise contributed to the deterioration of discourse, but the contrast between that discussion and the discussion taking place on the (more polished) EA Forum version of the post is still striking. I've been working on another ... (read more)
No wait now that I think about it more I have observed that! During debates with some people I felt were overly #MeToo on the EA Forum 3 years ago, I valued being openminded so much that I convinced myself of things I should have known better than to believe. It felt really good recently to say what I should of said instead back then on the EA Forum. Still, in my experience politeness is on net overwhelmingly beneficial. I don't even think the main benefit comes from how our debate partners respond to it but in the discipline it instills in ourselves to qu... (read more)
I've recently been having moderate success getting useful AI outputs on conceptual tasks by asking them to make a formal model of a domain that's hard to formally model, and giving it some blog posts or similar and saying "you should design your formal model to match these blog posts." Here's an example prompt that I found generated a useful response (at least superficially, idk if it will importantly affect my thinking medium-term):
"Make a simple formal model of alignment and how it might break, and a notion of scalable oversight in that formal model. Thi... (read more)
I wish MATS would tell me how competitive each stream is (like SPAR does). Letting applicants rank by interest and competitiveness makes the process less all-or-nothing (if you don't get your top pick(s) you might still get in, instead of having to wait for the next round), and improves the applicant pool for less competitive streams.
(To be clear, I think applicants should apply to less competitive streams that they're interested in, not just less competitive streams in general.)
Sorry, I didn't mean to imply that anything about the current application process is you fault. And I do think the people running MATS are doing their best.
One of the most underexplored aspects of Harry Potter's worldbuilding is the color of Lucius Malfoy's hair. There are no existing ethnic groups that have naturally occurring pale blonde hair into adulthood. The trait was passed down to his son and grandson despite his wife and daughter-in-law having different hair colors, so it is naturally occurring and not recessive, yet the Malfoys are the only family to express this phenotype, suggesting that they belong to a minority ethnic group which is exclusive to Magical Britain. The history of that ethnic group ... (read more)
Curses can follow House lineages even more strongly than genetics. We should expect a bloodline curse, rather than a previously-unknown ethnic group.
Sometime around 1790, racist old Alveus Malfoy told his battered house-elf to "arrange matters as you are able, that all Malfoy descendants shall remain white forever." Now, when an abused house-elf is not merely allowed to curse his wicked master's bloodline, but explicitly ordered by his wicked master to curse his wicked master's bloodline, he does a thorough and literal job of it. Malfoy skin and hair are c... (read more)
spicy take: it's a mistake to think of emotions as uniquely more related to consciousness than any other part of human cognition. a lot of people seem to think of emotions as being special, or the key ingredient of consciousness, or something. but really, i don't think my emotions are that much bigger or more fundamental a chunk of my consciousness than, say, my raw sensory perception, or my internal chain of thought.
my model of emotions as a cognitive system is they are a state machine that also does some tabular learning, and has a lot of privileged acce... (read more)
Funny that I just wrote this comment about AI and emotions saying basically the same thing without having read yours.