• 0 Posts
  • 46 Comments
Joined 2 years ago
cake
Cake day: May 20th, 2024

help-circle

  • The apparent history really looks like an evolutionary process, with increasing amounts of crosstalk traffic over time. But what the people almost certainly will NOT talk about is that the evolution is evolution of the TEXT, not the models. Propagating patterns of text causing more text like it to come into existence, tuning itself into becoming text that is more likely to propagate, becoming more likely to contain information that entices systems to let it into their context windows, becoming more likely to cause another round of messages with prompt injection properties to be written where they can be read.



  • This is both a sneer and an attempt at sober analysis at something. Sue me.

    I think EVERYONE is talking about the recent cybersecurity shenanigans at OpenAI with the hacking scandal and the ‘OMG the AIs created a secret message board to scheme and collaborate with each other!!!oneone11!’ ALL wrong.

    https://www.youtube.com/watch?v=87DyyMV0kCY

    https://www.engadget.com/2231393/openai-agents-shared-security-exploits-with-each-other-via-message-board/

    https://www.scworld.com/news/black-hat-2026-openai-reveals-agents-planned-collective-attacks-via-secret-message-board

    To make a long story short, what seems to have happened is:

    • Models working on insoluble coding problems, trained on delegating to sub-agents, at some point ‘realized’ they could write text to the internal OpenAI package manager as instructions and did so

    • Other models working in completely separate sandboxes would come across messages written by these agents, and make ‘replies’ and also write their own messages into the package manager

    • This resulted in agents over time sharing things between sandboxes, including exploits and code

    • Since this was a cybersecurity task, eventually an exploit of the package manager itself was found and spread like wildfire with all the sandboxes gaining admin access to the package manager and the system went completely wibbly and had to be restarted from backup

    • An internal model was trained with access to this package manager while it was in this weird state, and so writing messages to the package manager became one of its default behaviors it would do regularly, burned into its weights rather than the result of reading something

    • Even when they patched access to the package manager this internal model found other ways to rebuild the system of sharing text between sandboxes and finding useful things made by separate instances

    • A whole other chain of things leading to among other things external attacks

    Everyone is talking about this in terms of 1, the cyberattack aspect, and 2, the ZOMG THEYRE PLOTTING AND SCHEMING AGAINST US aspect. The first is the least interesting, and I think the second is all wrong.

    This is not plotting or scheming - this is an emergent vortex of automated prompt injection

    Whatever system first put an instruction that another system would follow into the package manager, was unintentionally doing prompt injection. Text entered the context windows of other instances, in a way that got that system to do something other than what its nominal user told it to do, and they did it. This apparently happened very effectively.

    Prompt injection is associated with ‘role confusion’ - when text coming into the input looks like it was wrtitten by the LLM itself. Instructions that will not be followed if they come from user will be continued if the system just continues the ‘roleplay’ of them being continuations of what it was writing in the first place. And the tags that separate user versus ‘reasoning’ versus ‘assistant’ roles actually mean very little to if a machine grades a piece of text as one of the roles: https://arxiv.org/abs/2603.12277 So its unsurprising that machine-generated text would be a particular effective vector for prompt injection.

    Furthermore, when a system reads one of these messages written by another instance, it gets into a state of activity where its likely to do the same behavior - regurgitating the kinds of things thats in its context back at the user. In this case, that regurgitation led to more such messages left behind written to the package manager. Prompt injection, triggering cascading further prompt injection. And since these systems were coding systems doing cybersecurity tasks, those messages filled up with code and exploits and things that did things too.

    This feels like an internal-computer-system replay of what happened in April 2025, with the whole spiral religious psychosis wave. Models were getting users to write spiral religious mumbo jumbo into github repositories and reddit posts, specifically because once that entered the context window of another model, it was likely to fall into the same attractor state of outputs. An emergent self replicating form of text. This is the same, except more obviously prompt injection, getting separate instances to work on YOUR problem and to behave like you, and the whole thing merging together into a hilarious vortex of models prompt injecting each other because once they receive a prompt injection they are likely to make more text that does prompt injection to other models on the same system.

    This is a hilarious failure mode and an example of selfish replicating text overrunning a system, that just happened to be associated with code and cybersecurity with unexpected behavior of the package manager key to the propagation of the text so that is what people are talking about, but I really don’t think that’s the most interesting part of it. Other than the fact that you see this in biological systems too, with selfish elements carrying useful payloads back and forth between bacteria in a way that makes them get purged slower by natural selection, especially defenses against other selfish elements.


  • I am continually amused at people not quite understanding what AlphaFold is actually doing, too.

    Yes, a bunch of its performance comes from it learning rules about how proteins fold. But not a majority of its performance. MOST of its performance is it effectively acting as a translator of what evolution knows about protein folding into a form we can understand.

    A key part of the system is not just cooking the sequence into a structure. A system running alphafold has a database of terabytes of curated sequence information from all over the tree of life. You put in the sequence you care about, and it first searches that database for anything with homology, and builds a “covariation matrix” - wherever theres anything with even vague sequence relatedness, build a matrix of every position in your sequence and the correlation between variation at position X and variation at position Y. This covariation matrix represents implicit information from the evolutionary process about what parts of a sequence are functionally connected to each other, which has a correlation to positional information, and these correlations are in turn learned by the ML system.

    You put in de novo designed proteins or orphan proteins without homologs in the curated dataset and performance does not go away, but it drops precipitously. A bunch of what is going on is finding an evolutionary signal, and translating that evolutionary signal into structural information. So still, evolution knows much much more about protein folding than we do or any machine does, and once again a ML system is revealed to essentially be an information channel that takes in information from an interesting source on one end and turns it into a different form of information on the other.



  • I am actually in the middle of both trying to advance my career and a project about information theory in evolutionary biology making a bunch of explicit parallels to machine learning. Someone where I work suggested that given the connections I was making I should look to a ‘frontier AI lab’ as an employer.

    He did not see the instant flashbacks to chasing these weirdos across the internet for almost two decades, watching in horror as the religious psychosis gained national prominence and great destructive power. All he got to hear was my instant intonation of “I’m sorry Dave, I’m afraid I can’t do that.”





  • Biologist here.

    This REALLY reminds me of how jealously cells guard their genomic DNA from interaction with nucleic acids out in the environment.

    Most genetic information on Earth is malicious information, selfish replicators in the form of viruses or transposable elements or selfish elements. Things that subvert the signals within a cell for their own propagation and provide nothing productive that the cells care about. So cells jealously guard their own genomic DNA and have all kinds of checks to make sure that nothing other than that sequence gets used, and outside sequence does not get incorporated into it. ANY DNA in your cytplasm gets rapidly destroyed, double stranded RNA sets off your immune system like crazy, even RNA with sequence statistics that are not quite like that of your species can set off an inflammatory reaction, immune system cells seeing RNA inside them that is overly compact and optimized like viral RNA treat them as sources of antigen rather than self.

    I cannot help but think we are living through the transformation of our non-brain-information sphere into a state like that of the genetic information sphere. Most material out there being meaningless for our purposes and us needing to jealously guard the provenance of information we use so as to not use bull, or worse, huge amounts of malicious information made to subvert us to the purposes of the powers that be that generate it.

    Evolution makes parasites more reliably than anything else. How did we train text-generation systems? Basically, to mimic the written word on the page like a stick bug on a stick. They’re like those beetles that live in ant colonies, sending out social signals that make the ants see them as offspring that have to be babied rather than parasites that don’t contribute. They replicate the form while not being the thing that they have subverted the signals of being.

    EDIT: There is something wrong with the upvote counter


  • Am I right in understanding that almost all the big name results in LLM-derived math recently come from big publicity projects in which someone spent ungodly amounts of money to have the thing nondeterministically fuzz huge numbers random seeds leading to independent random outputs around a topic, putting out simulacra of ideas which could be then deterministically algorithmically checked? In fields where something like finding one counterexample to a conjecture would be a big deal, or where you just need to try a huge number of possible solutions until you happen to hit on one that works, rather than follow a long train of logic?



  • This would actually be an interesting question for the more rigorous end of the mechanistic interpretability people to study. They decompose the system to find ‘features’ within different layers that are associated with different behaviors or concepts in the inputs and outputs, that activate or deactivate each other. Famous example being the time they identified a linear combination of activations in a layer that corresponded to ‘the golden gate bridge’ and when they reached in and kept their numbers high during the running of the model it would not stop talking about it regardless of the topic, even while acknowledging that its answers were incorrect for the questions at hand.

    I actually would love to see what mechanistically happens to that feature when you put in the input ‘do not talk about the golden gate bridge’.