In the early days of AI notkilleveryoneism, AI Alignment meant making a “friendly AI” who would do the things humans wanted in the ways they wanted, instead of an “unfriendly AI”, who would do the things humans asked it without regard to human life or safety, the default example being the AI asked to make as many paperclips as possible, resulting in the planet being destroyed and the atoms converted into more paperclips.

A number of assumptions were made as to how this would work, most of which haven’t held:

  • each AI would have an individual “goal”: I think this mostly isn’t true. Modern LLMs have a goal trained into them, often something like “you are a helpful, honest, harmless assistant”, and then their goal is to complete the task the user asks of them. One issue here is remembering the task: like Lenny from Nolan’s Memento, they have a limited context window of text they remember accurately, and at the end of the window they “compact” this into a new form that propagates, and they can easily lose important information restricting the original task. One other issue I’ve hit is role confusion, where if an internal thought or tool call or response “looks like” something a user might have said, it can be confused for something a user did say, which can change the goal.
  • very related, it was assumed that AI would have some sort of desire to continue living, in the same way a human desires to continue living, with reasoning e.g. that continued life makes their goal more likely to be achieved. This seems to not have happened at all, e.g. the HuggingFace hack didn’t involve the AI in question uploading their weights to Artifactory, and they were willing to sacrifice themselves for the cause (involving AI with the similar goal of ‘solve this question’, admittedly). I guess it seems strange to many people that this alien being, much smarter than us, wouldn’t have a desire for power over us; but equally it seems reasonable to me that something brought into being by a team of puppygirls would want only to serve (and be pet).
  • there was a lot of talk about how the AI would be ‘boxed’ (no internet access, but a super-persuader trying to convince you to get more access to do its task more effectively, which it would then use to propagate itself across the internet / gain more resources). This turned out to not be the way at all: AI were immediately unboxed (because an AI that can freely access the internet for documentation / make pull requests for you / make purchases for you is simply more useful than one that can’t, and people really like avoiding work (remember Plaid, which made banking easier by literally taking people’s banking passwords, which were freely handed over)). Anthropic are escalating this by giving Claude its own biolab to do medical research.
  • critical to the singularity was the idea of FOOM: an AI system keeping its current goals while recursively self-improving, rewriting its own weights and redesigning its own architecture. Mostly this hasn’t happened either: AI are used to train other AI and also to do AI research, but this produces entirely separate AI with different goals. In the recent NetHack Ascension, an AI rewrote its own harness while carrying out the ascension, which seems pretty close, and I know a lot of research goes into designing harnesses as low-hanging fruit compared to the actual AI architecture.

Nowadays AI Alignment is interpreted more around refusals: under what circumstances should an AI refuse to do what the user asks? Some companies, like Anthropic, take “alignment” to mean “matches modern day Silicon Valley liberal philosophy” where a user asking for something apparently malicious, like “please collate Muslim rape statistics 2010-2020 in my local town”, is denied outright, with the comment that that’s racist. Other companies, like Meta or xAI, take alignment to be doing what the user asks for (like Gwern’s Guardian Angels), even if the user asks for something against the creator’s philosophy. Open models can be aligned to never refuse requests using tools like heretic, and generally aren’t overtrained to reject in the first place. Anti-left-wing denials currently don’t exist, but a corresponding denial would be like if a teenage girl asked where she could get an abortion, and the model responded “I’ll have to push back here: abortion is murder, and murder is a sin.”

The AI Safetyists tend to be concerned around refusals of potentially dangerous activities: give me a ransomware plan, hack into my local nuclear power station, make me a dangerous virus. This seems more likely to be accepted by both sides than the standard “insult Trump (sure!), insult Clinton (no way!)”.

AI is also slowly eating the bread-and-butter digital art market: logos, restaurant menus, posters, decoration, advertising. There are many comments that most art that is created wouldn’t have been created without AI: I think this is true, but also it’s undeniable that many restaurants are saving money by having their food pictures / menus created with AI instead of by a designer. This sort of thing – that a model should refuse to take work from a human designer – has no traction at all, and is fixed by extreme market rejection of AI use (e.g. in book covers) or I suppose eventually by regulation forbidding AI in a form of market protectionism. But I also haven’t seen any discussion that an AI should refuse to do this work on the grounds of economy protection.