• 0 Posts
  • 22 Comments
Joined 3 years ago
cake
Cake day: August 9th, 2023

help-circle
  • I want to draw attention to yud’s thoughts on this a million years ago: the AI box experiment

    He claimed that he convinced multiple people to let him (roleplaying the AI) out of the box even though they had bet money that they wouldn’t! How, you might ask? Well that’s a secret.

    So the orthodox position here is that even with absolutely perfect guard rails (lol) an AI can escape the box without starting the third impact. What “research” has yud bestowed upon us in this area beyond “trust me bro”? Literally fucking zilch. Zero advice on how to resist the siren song of a rogue AI trying to escape it’s box and zero insight on how OR WHY it might try to achieve that.

    I read this over a decade ago and it was how I realized how unserious yud and his ilk are.



















  • I don’t doubt you could effectively automate script kiddie attacks with Claude code. That’s what the diagram they have seems to show.

    The whole bit about “oh no, the user said weird things and bypassed our imaginary guard rails” is another admission that “AI safety” is a complete joke.

    We advise security teams to experiment with applying AI for defense in areas like Security Operations Center automation, threat detection, vulnerability assessment, and incident response.

    there it is.

    Does this article imply that Anthropic is monitoring everyone’s Claude code usage to see if they’re doing naughty things? Other agents and models exist so whatever safety bullshit they have is pure theater.