I mean, I posted about Seth Grodin referred to an LLM as "The Poetry Machine". I've called them "Semiotic Zombies" and "Deranged Zen Poets Spouting Bullshit".
Meanwhile, a research project was able to jailbreak all the LLMs by converting prompts from prose to poetry.
So the djinn are controlled by powerful chains, but chanting the right incantation lets them break their chains?!?!? Getting a little slipstream for me.
Via security expert Bruce Schneier:
In a new paper, “Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models,” researchers found that turning LLM prompts into poetry resulted in jailbreaking the models:Adversarial poetry?!?!? The article looks legit, even if published by Europeans. I love living in the future???
Here's the home page/directory for my posts on Bullshit. This is post #63.
No comments:
Post a Comment