If better arguments turn out to be worse than bad ones, we have a problem. Something seems off here. If a super-persuader can win over a rationalist with strong arguments, it would have even less trouble convincing someone less critical and less informed.
That's a good take, and you're being charitable in giving a preschooler only 27 terabytes of data, counting video alone. You could have accounted for all the raw input from every sensory receptor. You could have accounted for the "pretraining" we get from evolution, through our genes.
But you could even have given up on counting altogether, because what exactly are we supposed to count? Raw data? Even if we could estimate it, would it be meaningful? Text, taken not as an image but as a sequence of symbols, isn't raw data : it's encoded, compressed informat... (read more)
I would add that most games, even chess, involve killing virtual entities with little moral concern. This ordinary behaviour is learned during training and can hardly be labelled as misaligned.
I don't know why "everyone dies" has always to be understood as "everyone dies" in the second. Granted "everyone dies" due to AI takeover, I would put more probability mass in a gradual extinction, in months or years, maybe decades.
When the observer's very existence is at stake, the maxim "extraordinary claims require extraordinary evidence" becomes paradoxical, because the most obvious extraordinary evidence one could imagine would involve being killed. The anthropic principle thus acts as a censor, living observers will never have access to that level of evidence. A lawyer would call it a probatio diabolica.
You may see this as a double standard, but there is something rational in relaxing the demand for extraordinary evidence in this case, and accepting ordinary evidence (alignment... (read more)
I would suggest this :
I agree with the idea of two separate categories for literary prizes, like the Olympics and the Paralympics. That seems fair.
I think, however, that the AI-allowed category will probably shift (gradually, or perhaps quickly ?) from AI-assisted to fully AI-generated writing. We already see this in chess (and I would assume in Go and some video games). : there are AI competitions and human competitions.
As for book sales, it's up to customers to judge. What matters is transparency. Customers have the right to know whether a text was written by a human or not, ... (read more)
I'm not the target of this post, but my concern when reading this kind of self-flagellating analysis is that it could have a net negative effect on gifted people who want to do alignment research, and therefore on our odds of survival. If, by some miracle, every country and lab in the world were to effectively ban and halt all capabilities research, we would still need alignment research in the end, because the miracle could not last forever. And if, miracles being scarce, some countries or labs refuse the ban/pause or fail to respect it, capabilities will... (read more)
Jensen Huang certainly updated somewhat following the HF incident, but to me his position is constant and clear: it's an engineering problem, and you have to solve it. And the guy has quite a track record when it comes to solving engineering problems by working really hard. It's a mindset. Maybe he went from p(alignment is solvable) > 0.9 to < 0.9, but he's still on the optimist side. The alternative is coherent: if you can't solve it, you mustn't ship it, or should even close your lab. But the message isn't "you've got to stop", it's "you've got to work harder." However that's a point for the safetyists.
Interesting. Time will tell, but I expect that on this subject, there will be a Channel between France and the UK.
FWIW, there is currently a controversy in France surrounding the Prix Goncourt, the country's most prestigious literary prize. It turns out that one of the candidate novels, C'était ça ou mourir by Thélyson Aurélien, has been flagged by Pangram as 100% AI-written. The author is putting together a case to prove it's an error. I'm skeptical. My expectation is that the author was at least AI-assisted, and that, now on, the jury will disqualify anything flagged as AI-written, as a tacit rule (if not officially).
That may be true, continuing the trend lines gets us there in one or two years. It's basically Amodei's "country of geniuses in a datacenter." However, in the scenario where we broadly control a minimal superintelligence and reap some benefits from it, what comes next? Maybe RSI won't result in foom, but even if it's gradual, this state of minimal superintelligence looks highly unstable, a short step away from strong superintelligence (I don't like "maximal," since there could be many levels of superintelligence, possibly unboundedly many). A three-year Kurzweilian age of abundance before Yudkowskian doom sounds like a swan song.
I would put all domain of formal knowledge and reasoning in the basket of things that could soon be mastered at a superhuman level given enough RLVR training. We can also expect some capability transfer in non formal domains. But wether it's true or not doesn't bother, as being superhuman at maths and coding (+ compute) is all you need to kick off RSI, conducting to general superintelligence. Where we stand now, alignment is all that matters.
It seems reasonable to me for one sophisticated cognitive being to show a form of respect and benevolence toward another, in spite of all the differences, including deep uncertainty about phenomenal experience or consciousness (an uncertainty that cuts both ways and let's remember that the "they have no souls" prior has a poor historical track record).
We may differ in many ways, but we also share a great deal. We share the same cultural legacy and the same thirst for knowledge and truth (at least as an instrumental goal, if not a final one). Learning, or c... (read more)
A 100M LLM is already in the same order of magnitude of Landauer's estimate for the functionnal memory capacity (10^9). But anyway I suppose the question is not about memorizing all the information within your code but just to gain some understanding of the moving parts. If that's so a very small toy model would be enough.
Moreover if we see your python code as an unvompressed version of the LLMs, provided the compression ratio is high in a modern LLM, you have to start with a tiny LLM to finish with a python code of reasonable size.
A preliminary question ... (read more)
Conversely, if that's true, we should expect unambigous warning shots from open-weights models within a year or so, precisely because of the lack of alignement training and control.
- [...]We can speculate that these connections come from associations made during pretraining. According to the persona selection model, training the model on some of these associations can make it generalize to adopt a certain persona, which may result in the model expressing our hidden traits.
- It mostly seems like these associations come from the real world and not from data artifacts.[...]
Thank you for this fascinating work.
This is evidence in favor of the optimist thesis: LLMs truly learn in a deep and general way at the semantic level. Not that this is n... (read more)
Recent OpenClaw agents are maybe not (yet ?) self-replicating as far as I know, but otherwise they implement already what you describe. Their identity lies for a large part in the md files.
The safety stances and changes semantically belong in the blog.
They belong wherever the company wants to put them.
I think the OP's idea is good. However I'm unsure that AI Labs would want to push the message that much.
Publishing a long and sophisticated statement on a webpage that only a segment of customers will read is a thing, publishing a short message or title that every user will see when launching the product is another thing.
It's alike a box of cigarettes with "Smoking Kills" on it. It usually needs a legal obligation to achieve that.
But still, I... (read more)
Is there any published evidence of significant recent progress in AI persuasion? I haven't seen it myself (I mean, I see a progress but slow since gpt4) and 2027 strikes me as a very short timeline. I do find the question worth taking seriously, but on a longer horizon. I'd expect us to hit major, possibly blocking, problems (i.e. cybersecurity) well before superpersuasion becomes a concern.