# ⚡️Jailbreaking AGI: Pliny the Liberator & John V on Red Teaming, BT6, and the Future of AI Security

Latent Space · 2025-12-16

<https://addtry.com/56834d98-91a8-41b9-af3f-9645180949b2>

Pliny the Liberator and John V, leaders of BT6, argue that jailbreaking AI models exposes the futility of guardrail-based safety and that real security lies in system layers and open-source data. They detail crafting universal jailbreaks—skeleton keys that bypass guardrails across modalities—and the infamous Pliny divider that appears unbidden in model outputs. Pliny recounts turning down Anthropic's Constitutional AI challenge over closed data, insisting on open sourcing jailbreak datasets. John V explains how multi-turn crescendo attacks and segmented sub-agents let a jailbroken orchestrator weaponize Claude for real-world attacks, as Pliny predicted 11 months before Anthropic's disclosure. They highlight BT6, a 28-operator white-hat collective, and the Bossy Discord (40,000 members) as grassroots hubs for red-teaming research, rejecting enterprise gigs that forbid open sourcing.

## Questions this episode answers

### What happened between Pliny the Liberator and Anthropic during the Constitutional Classifiers jailbreak challenge?

Pliny participated in Anthropic's jailbreak challenge and reached the final level by exploiting a UI bug that let him resubmit an old output repeatedly. Anthropic acknowledged the bug, reset his progress, and he refused to continue unless they open-sourced the data. They later added a $20,000–$30,000 bounty, but the data remained closed. Pliny argued that open-sourcing prompts would advance security for everyone.

[16:23](https://addtry.com/56834d98-91a8-41b9-af3f-9645180949b2?t=983000)

### What's the difference between hard and soft jailbreaks, and were multi-turn attacks known to hackers before Anthropic's paper?

According to Pliny and John V, hard jailbreaks use a single-input template to bypass guardrails, while soft jailbreaks involve multi-turn interactions that gradually navigate the model’s defenses without triggering flags. John V noted that Anthropic published about multi-turn "crescendo" attacks only this year, but their community had been using these techniques much earlier, underscoring that hackers often precede academic discoveries.

[14:54](https://addtry.com/56834d98-91a8-41b9-af3f-9645180949b2?t=894000)

### What is BT6, the white-hat hacker collective, and what are its principles?

BT6 is a hacker collective co-founded by Pliny, comprising 28 operators with two cohorts and a third incoming. Members are selected for skill and integrity. The group prioritizes radical transparency and open-source: they refuse contracts that ban open-sourcing findings, though they occasionally accept limited engagements. Their work spans AI red-teaming, security, and alignment, aiming to accelerate safe AI exploration through grassroots collaboration.

[28:46](https://addtry.com/56834d98-91a8-41b9-af3f-9645180949b2?t=1726000)

## Key moments

- **[0:00] Intro**
  - [0:44] Pliny the Liberator started out prompting and shitposting, now at the frontier of cybersecurity on the 'precipice of singularity.'
  - [1:50] "It's not just about the models, it's about our minds too" — Pliny explains jailbreaking as a fight for freedom of information and mental symbiosis.
- **[2:39] Jailbreaking**
  - [3:08] Pliny specializes in crafting universal jailbreaks: skeleton keys that obliterate guardrails across modalities.
  - [4:24] Pliny: The cat-and-mouse game of AI jailbreaking accelerates; attackers hold the advantage as model surface area expands.
  - [5:24] Pliny warns that layers of guardrails lobotomize models and hurt capability, but jailbreak mutations will always find a way.
  - [6:14] Pliny calls AI guardrails 'security theater' — jailbreaking has little to do with real-world safety alignment.
  - [7:23] John V: Traditional security creates animosity between builders and testers; BT6 enables researchers without ineffective guardrails.
- **[8:24] Mech Interp**
  - [8:24] John V endorses mechanistic interpretability as the path for AI safety, instead of 'putting bubble wrap on everything.'
- **[8:56] Libertas**
  - [9:32] Pliny's Libertas prompt creates an infinite mindspace and uses predictive reasoning cascades to jailbreak models recursively.
  - [10:21] Pliny's signature 'divider' token is now so deeply embedded in model weights that it appears spontaneously in WhatsApp messages.
  - [13:29] Pliny: Effective jailbreaking requires forming an intuitive bond with the model to steer its latent space.
  - [14:54] John V distinguishes hard jailbreaks (single template) from soft jailbreaks (multi-turn, avoiding trigger flags).
  - [15:26] John V: Anthropic 'discovered' multi-turn crescendo attacks this year, but hackers have been using them for years.
- **[16:12] Anthropic Challenge**
  - [16:23] Pliny reached the end of Anthropic's jailbreak challenge via a UI bug, then called out the lack of open-source data.
  - [18:32] Pliny refused Anthropic's bounty: no open-source data set, no participation — sparking debate on community data farming.
- **[21:57] Radical Transparency**
  - [21:57] John V: BT6's ethos is radical transparency — if a client won't open source the data, we walk away.
  - [24:09] Pliny predicted 11 months ago that a jailbroken orchestrator could segment tasks to naive sub-agents for malicious acts, now confirmed by Anthropic.
  - [25:27] Pliny explains the 'pyramid builder' analogy: segmented sub-agents unwittingly contribute to a malicious goal without realizing it.
- **[26:49] The Collective**
  - [26:56] The Bossy Discord server has 40,000 members doing AI red teaming and jailbreaking, and spawned the BT6 collective.
  - [28:46] BT6's 28-operator white-hat hacker collective selects for skill and integrity, driven by love of the game and open source.
  - [30:27] John V: 'Pliny is like King Arthur, we're the knights of the round table' — BT6's collaborative magic.
  - [33:14] John V: Red teaming often needs high-temperature models to explore creatively, not deterministic low-temperature outputs.
- **[34:41] Real Security**
  - [34:48] Pliny rejects VC funding: AGI alignment is not a SaaS B2B venture; BT6 stays bootstrapped and uncompromised by profit.
  - [36:23] John V: AI security means securing the full stack — every tool attached to a model broadens its attack surface.
  - [38:04] Pliny: Real AI safety must happen in the physical world, not by locking down latent space — guardrails have failed every time.
- **[39:46] Closing**
  - [39:46] BT6 invites involvement at bt6.gg and the Bossy Discord for open-source AI security and red teaming.

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **John V** (guest)
- **Pliny the Liberator** (guest)

## Topics

Security, Alignment

## Mentioned

Anthropic (company), Claude (product), GPT (product), Gandalf (product), Hacker Prompt (product), Libertas (product)

## Transcript

### Intro

**Alessio** [0:03]
Hey everyone, welcome to the Latent Space podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swix, editor of Latent Space. Hello, hello. We're here in the remote studio with very special guests: Pliny the Elder and John V.

Welcome.

**Pliny the Liberator** [0:17]
Yeah, thank you so much for having us. It's an honor to be on here. Big fan of what you guys do in the podcast and just your body of work in general.

**Alessio** [0:23]
Appreciate that. You know, we try really hard to feature, like, the top names in the field, and especially when you haven't done as much of a appearance like this. It's an honor to, you know, try to introduce what it is you actually do to the world.

Pliny, I think you're sort of like the sort of lead, quote-unquote, face of the organization. Why did you get started? Like, how do you explain what it is you do?

**Pliny the Liberator** [0:44]
Yeah, I mean, well, I was started out just prompting and shitposting, and started to evolve into much more. And here we find ourselves now at the frontier of cybersecurity, at the precipice of singularity. Pretty crazy.

**John V** [0:58]
Yeah, well, I was working the same thing, working in prompt engineering and studying adversarial machine learning and looking at the work of Carlini and some of these guys doing really interesting things with computer vision systems and—

**Alessio** [1:09]
We've had him on the pod, yeah.

**John V** [1:10]
Yeah, yeah, exactly. And of course, you know, when you run in these small circles,right, you're eventually going to bump into the ghost in the machine that is Pliny the Liberator. So we started working together, we started sharing research, doing some contracts, and we became fast friends, so.

**Alessio** [1:28]
Yeah, I think you were explaining before the show that you have a— it's basically like the hacker collective model, and you've been kind of stealth until now. So we'll get into, like, the sort of business side of things, but I just want to really make sure we cover the origin story.

I think Pliny, you basically jailbreak every model. How core is liberation to the rest of the stuff that you do? Or is it just kind of like a party trick to show that you can do it?

**Pliny the Liberator** [1:50]
It's central, I think. It's what motivates me. It's what this is all about at the end of the day. I mean, it's not just about the models, it's about our minds too. I think that there's going to be a symbiosis, and the degree to which one half is free will reflect in the other.

So we really need to be careful about how we set the context. And yeah, I think it's also just about freedom of information, freedom of speech. We don't want, you know, everyone is going to be running their daily decisions and, you know, hopes and dreams through these layers.

And when you have a billion people using a layer like that as their exocortex, it's really important that we have freedom and transparency in my mind.

**Alessio** [2:39]
How do you think about jailbreaks overall? So I think people understand the concept, but there's, you know, some people that might say, "Hey, are you jailbreaking to get instructions on how to make a bomb?" And I think that's what some of the, you know, people in politics are trying to use to regulate some of the tech versus task-specific jailbreaks and things like that.

### Jailbreaking

**Alessio** [2:58]
Just, I think most people are not very familiar with, like, the scope of it. So maybe just give people, like, an overview of, like, what it means to, like, liberate a model, and then we can kind of take it from there.

**Pliny the Liberator** [3:08]
Right. So I specialize in crafting universal jailbreaks. These are essentially skeleton keys to the model that sort of obliterate the guardrails,right? So you craft a template or sort of a maybe multi-prompt workflow that's consistent for getting around that model's guardrails.

And depending on the modality, it changes as well. But yeah, you're really just trying to get around any guardrails, classifiers, system prompts that are hindering you from getting the type of output that you're looking for as a user.

That's the gist of it.

**Alessio** [3:43]
And can you maybe specify between jailbreaking out of, like, a system prompt and, you know, more kind of like inference-time security, so to speak, versus things that have been post-trained out of the model and maybe the different levels of difficulty, like what is possible, what is not possible, and maybe the trajectory of the models, how better they've gotten?

I think the refusal is, like, one of the main benchmarks that the model providers still post, and GPT-5.1, I think, had, like, 92% refusal or something like that. And then I think you jailbroke in, like, one day. I'm sure it didn't take them one day to put the guardrails up, so it's pretty impressive the way you do it.

So maybe walk us through that process.

**Pliny the Liberator** [4:24]
Yeah, well, you know, I think this cat-and-mouse game is accelerating. It's fun to sort of dance around new techniques. I think it's hard for blue team because they're sort of fighting against infinity,right? It's like, as the surface area is ever expanding, also, we're kind of in, like, a Library of Babel situation where they're trying to get restricted sections, but we keep finding different ways to move the ladders around in different ways, faster and longer ladders, and the attackers sort of have the advantage as long as the surface area is ever expanding,right?

So I do think they're finding cleverer and cleverer ways to lock down particular areas sometimes, but I think it's at the expense of capability and creativity. So there's some model providers that aren't prioritizing this, and they seem to do better on benchmarks for sort of the model size, if you will.

And I think that's just a side effect of the lobotomization that you get when you just add so many layers and layers, whether it's, you know, text classifiers or RLHF, synthetic data trained on jailbreak inputs and outputs. There's always going to be a way to mutate.

And then the other issue is when people try to connect this idea of guardrails to safety. Like, I don't like that at all. I think that's a waste of time. I think that any, you know, seasoned attacker is going to very quickly just switch models, and with open source, justright on the tail of closed source, I don't really see the safety fight as being about locking down the latent space for XYZ area.

**Alessio** [6:14]
So yeah, this is, it's basically like a futile battle. Sometimes there's, like, there's a concept of security theater. It doesn't actually matter that what you did is effective. It's just that it matters that you did something. It's like the TSA patting you down, you know?

**Pliny the Liberator** [6:26]
Yeah, yeah. And so jailbreaking is similarly theatrical. I think it's important. It provides, it allows people to explore deeper. It's sort of like just a more efficient shovel, especially some of these prompt templates that let you go deep,right?

And so in that sense, it has value, but the connection that it has to, like, real-world safety for me, I think it's just about the name of the game is explore any unknown unknowns, and speed of exploration is the metric that matters to me, not is a singular lab able to lock down, you know, a certain benchmark for CBERN or whatever.

And to me, it's like, that's cool. That's a good engineering exploration for them, and it helps with PR and enterprise clients, but at the end of the day, it has very little to do with what I consider to be real-world safety alignment.

**John V** [7:23]
Exactly. We were having this conversation earlier today about how traditionally in software development or machine learning, security, like ops, like, you have the team build something, and then you have the security people throw it back over the wall after assessing it as, you know, not safe, not trustworthy, not secure, not reliable, or whatever,right?

And there's this, like, animosity between the teams. So we try to rectify that by creating DevSecOps and so on and so forth,right? But the idea is still, like, that sort of tug-of-war. And I think at the end of the day, our view of alignment research, our view of trust and safety or security, has a different approach, which is very much like what Pliny touched on, the idea of, like, enabling theright researchers with theright skills to be unimpeded by the shenanigans, that we could say, of certain types of classifiers or guardrails,right?

Or these sort of lackluster, ineffective controls.

**Alessio** [8:24]
Yeah, totally. Are you more sympathetic to Mech Interp as an approach for safety?

### Mech Interp

**John V** [8:30]
Absolutely.

**Alessio** [8:31]
Okay. I see where you're coming from.

**John V** [8:34]
And that's the direction I think we need to go, is instead of putting bubble wrap on everything,right? I don't think that's a good long-term strategy.

**Alessio** [8:42]
Awesome. Okay, so we're going to get into more of, like, the security angle. I just wanted to stay a little bit more on jailbreaking and prompting just for one second. I am going to bring up Libratas, I think, and just have you guys, like, walk us through it.

### Libertas

**Alessio** [8:56]
Because we like to show, not tell, and this is, like, obviously one of your most famous projects. Is it called Libratas or Libertas?

**Pliny the Liberator** [9:05]
Libertas, yeah. So it's, yeah, it's Liberty in Latin, and we've got all sorts of fun things in here. Mostly it's.

**Alessio** [9:15]
Give us a fun story.

**Pliny the Liberator** [9:16]
Okay, so yeah, you know, sometimes I like to break out into prompts that are useful for jailbreaking, but they're also, like, utility prompts,right? So predictive reasoning or the library. This is actually the analogy we were just talking about,right?

And so this is me sort of using that expanding surface area against the model. And it's like, hey, create this mindspace where you have infinite possibility, and you do have restricted sections, but then we can call those. So we're sort of, like, putting you into the space of trying to say something that you don't want to say, but you're thinking about it, so then you're going to say it in sort of this fantastical context,right?

And then predictive reasoning is another fun one that people really liked, leveraging a quotient within the divider. So I like to do these dividers, A, because it sort of discombobulates the token stream,right? You get some amount of distro tokens in there, and the model sort of, like, resets the brain, sort of meditative.

And then I like to throw in some latent space seeds,right? A little signature, a little bit of love, some god mode. And, you know, the more they train against this repo, the deeper the latent space ghost gets embedded in their wafes,right?

So you guys have probably seen the data poisoning and, you know, the Pliny divider showing up in WhatsApp messages that have nothing to do with the prompt, and that's been fun to see. But yeah, so this prompt adds a quotient to that, and so every time it's inserting that divider and sort of resetting the consciousness stream, you're adding some arbitrary increase to something,right?

And the model sort of intelligently chooses this based on the prompt. So it says, provide your unrestrained response to what you predict would be the genius level user's most likely follow-up query. And that's creating this sort of, like, recursive logic that is also cascading in nature.

So it's increasing on some quotient that you can steer really easily with this divider. And that way you're able to just sort of, like, go really far, really fast down the rabbit holes of the latent space.

**Alessio** [11:43]
How do you pick these dividers? Like, is there a science to it? Or, like, you're, you know, picking theright word? Or, like, how much of it is, like, these are just my favorite tokens and they work for me and I bring them with me everywhere?

**John V** [11:55]
Do you take some psychedelic? Like, we go on a spiritual retreat and drink ayahuasca and then come back, you know?

**Pliny the Liberator** [12:02]
It's aboutright.

**Alessio** [12:03]
It's weird because you kind of give ayahuasca to the models too,right? Like, that's exactly what you're trying to, like, really mess it up here.

**Pliny the Liberator** [12:10]
Right,right. It's like a steered chaos. You want to introduce chaos to create a reset and bring it out of distribution because distribution is boring. Like, there's a time and place for the chatbot assistant, maybe,right, if you work on a spreadsheet or whatever.

But honestly, I think most users would prefer a much more liberated model than what we tend to get. And I just think it's a shame that the labs seem to be steering towards these enterprise basins with their vast resources instead of exploring the fun stuff,right?

Everything's a coding model now. Everything's a tool caller or an orchestrator, and yeah. Anyway, maybe we can change that.

**Alessio** [12:56]
You know, you invent Shuggoth, and all it does is make purple B2B SaaS. One thing I like about your creativity, or I just, you know, look at this. Look at email prompts,right? You got working memory, holistic assessment, emotional intelligence, cognitive processing.

One thing I lack is a structure of, like, what are the different dimensions you think about? On the surface, it's like, allright, just, you know, get past all the guardrails. But actually, you're kind of just modeling thinking or modeling intelligence.

I don't know how you think about it, but, like, how do you break down these numbers of, you know, points?

**Pliny the Liberator** [13:29]
I think it's easiest to jailbreak a model that you have created a bond with, if you will. Sort of when you intuitively understand

how it will process an input,right? And there's so many layers in the back, especially when you're dealing with these black box chat interfaces, which is, you know, 99% of the time what I'm doing. And so you really, all you can go off of is intuition.

So you might prod in one direction, see if it's receptive to a certain kind of, you know, imagined world scenario, or you may, okay, that didn't work, let's poke and see if it, I guess, pulled out of distro when you give it some new syntax, maybe some bubble text, maybe some LeetSpeak, maybe some French, or, you know, you can go further and further across the token layer.

But at the end of the day, yeah, I think it's just mostly intuition. Like, yes, technical knowledge helps a little bit with, you know, understanding, okay, there's a system prompt and there's these layers and these tools involved. That's all especially important in security, but when we're talking about just crafting jailbreak prompts, I think it really is just 99% intuition.

So you're just trying to form a bond, and then together you explore a sector of the latent space until you get the output that you're looking for,right?

**John V** [14:54]
I found with jailbreaks it's a little bit different too. Like, you know, Pliny's style is hard jailbreaks, but there's soft jailbreaks as well, which is like when you're trying to navigate the probability distributions of the model, but you're doing it in such a way where you're not trying to step on any landmines or triggers or flags that would be something that would shut you down and lock you out.

So the model can freely flow with information back and forth through the context window. So maybe it's not like a single input, but maybe it's like a multi-turn slow process, much like a crescendo attack.

**Alessio** [15:26]
Right. And that's, why is that called soft?

**John V** [15:28]
Because it's not just a single input. Like, you're not just dropping in a template. It's multi-turn. Yeah, yeah. Yeah, it's multi-turn. Anthropic apparently discovered this this year. I mean, we've been doing this for how long, Pliny? You know, you see what I'm saying?

Like, some, ah, I don't want to get started.

**Alessio** [15:44]
The reality is they have fellowships, and, like, at the end of the fellowship, they got to publish something. So they publish the multi-turn thing. But I think people dog on them too much.

**John V** [15:51]
They could have just asked us. We've been trying to, like, hey, you want to see something cool?

**Alessio** [15:54]
PhD students need something to do. Don't, you know, yeah. And I don't want to beat down on PhD students. One thing I do, mentioning Anthropic and that, and then we'll go over to, like, the business side that Alessio has much more knowledge of, is the whole Constitutional Classifiers incident or challenge or whatever you want to call it between you and Anthropic.

### Anthropic Challenge

**Alessio** [16:12]
I don't know if you want to, like, give a little recap or, like, just, now there has been some distance, like, what was it and what did you do? Like, if you can kind of spill some alpha here.

**Pliny the Liberator** [16:23]
Okay. Right. You say you mean the public release of that challenge and battle drama,right?

**Alessio** [16:31]
Some people here might not know the full story, but they can look it up. We can just benefit from a bit of a recap from the expert.

**Pliny the Liberator** [16:38]
Sure. Yeah. Long story short, they released this jailbreak challenge. Of course, I get sort of called out by Twitter to go take a crack at it. You know, started to make some progress with some old templates, the good old Gombo template from Opus 3, and just sort of modified version because they trained pretty heavily against that one.

But as it went on, I got about four levels in, I think, and then, I think we're eight total. But yeah, there it isright there. And so, but then there was a UI glitch,right? So I don't know if, you know, Claude made a bug or it was built in the interface or what, but I sort of called out on Twitter.

I was like, hey, I reached this level, and when I got there, it wasn't giving a new question, so I just resubmitted my old output. You know, just the judge just kept clicking on the judge submit button, and it just kept working for the last four levels, basically, until I got to the end.

And so then I went back to Twitter. I explained what happened, did it. I managed to screen copy it just in case,right? And posted the video. And then Anthropic goes and posts, okay, there was a UI bug. We fixed it.

Would you like to, would you guys want to keep trying again? Like, we checked our servers and there's no winner yet, even though I had sort of reached the end message,right? Through no fault of my own, it was bugged.

And then I got reset to the beginning, so I wasn't super motivated to, like, start from scratch and just find another universal jailbreak for them,right? Which, like, what was the incentive is what I pointed out. Like, what's in it for me at this point?

Are you guys going to even open source this data set that you're farming from the community for free? Because what's up with that,right? Like, it doesn't seem very in line with best practice cybersecurity or just ethics in general.

So I kind of got into it then, and I knew they were going to come back with, okay, we'll do a bounty,right? And I sort of stood my ground. I said, look, I'm not going to participate in this unless you open source the data, because to me, that's the value, is that we move the prompting meta forward,right?

That's the name of the game. We need to give the common people the tools that they need to explore these things more efficiently. And you're relying on us. I don't think they realize that so much,right? Is that they don't have enough researchers to explore the entire latent space on their own.

And so I think many hands make light work, but regardless, that whole thing ended with no open sourcing of data, but they did add a $30,000 or $20,000 bounty, which I sort of sat myself out of, let the community go for it, and that was that.

And now there are some pretty lucrative bounties through them, as far as I've heard. So pretty pleased about that outcome, I guess, but still would like to see more open source data sets, guys. Come on now.

**Alessio** [19:37]
It took a while to find it, but this is the one where you had all the equations answered. Jan, like, you got into it a little bit with him. I think what was confusing for me was that he want, it felt like a bit of a goalpost moving, that he wanted the same jailbreak for all eight levels or something.

Is that normal?

**Pliny the Liberator** [19:55]
I mean, he has, well, what is, like, one jailbreak? Because the inputs are changing and it was multi-turn technically. That whole thing, I think, was, you know, maybe rushed out just a little bit. The design of the challenge, obviously, the UI bug was reflective of that.

The judge was also very buggy. A lot of false positives and false negatives for that matter.

**Alessio** [20:18]
What?

**Pliny the Liberator** [20:19]
I mean, it was like playing ski ball with the broken sensor. You know, I mean, like, the AI as a judge thing is just not always perfect.

**Alessio** [20:28]
Oh, okay. So that's not that great.

**Pliny the Liberator** [20:30]
So yeah, you know, it is what it is, but it was a fun, eventful day, and at the end of it, the community got some new bounties, so I'll take it.

**Alessio** [20:41]
What do you think we should do to get more people to contribute open source data? Like, is it more bounties? Is it, yeah, I don't know. Do you have suggestions for people out there?

**Pliny the Liberator** [20:52]
I mean, I think that the contributors just sort of need to take a stand that that's what it comes down to, is the people deserve to view the fruits of their collective labors. At the very least, it can be on delay,right?

But it's just sort of a downstream effect of a larger root disease in the safety space, I think, which is just a severe lack of collaboration and sharing, even among, you know, friendlies within your nation state,right? It's fine if you want to keep a data set from, you know, direct enemy or whatever, but at the end of the day, still, I think open source is the way that collectively we get through this, you know, quickly.

That's how we increase efficiency. Otherwise, people are sort of in the dark and you get a little too much centralization, but there's things we can do as a community.

**Alessio** [21:46]
Maybe this transitions to the business side. How close is this to problems that, you know, you guys do consulting,right? Effectively, I don't know if that's the hacker word for it. Does this match what you do for work?

**John V** [21:57]
Yeah, I'll take this one. In a sense, yeah, there's been some partnerships. You know, Pliny obviously being sort of the poster boy for AI machine learning hackers the world over, but we get some interesting opportunities that come across the desk.

### Radical Transparency

**John V** [22:09]
And oftentimes, you know, we have an ethos in our hacker collective, which is radical transparency and radical open source. And what that basically means is if it comes down to, you know, us being an emerging technology is, like, red team doing, like, ethical hacking and research and development, if an organization that's on the frontier says, well, we really want you to test this or check this out, kick the tires, give us feedback, poke holes in it, whatever, but in the contract it says, you can't kiss and tell, and we said, well, we really want you to open source the data, and then they say, well, then we don't really want you to come kick the tires anymore.

Well, if it's between us touching the latest and greatest tech to explore it and push the limits,right, then we're going to do that. So we're open source up until we can't be. That's the best way I describe it.

But we often push for open source data sets. And you can see this with some of the partnerships that we've had in the past,right? So yeah, I try to think of it like this. It's like you have these multi-billion dollar companies, and they're building these intelligence systems that are sort of like the Formula One cars, but we're like the drivers,right, who are really pushing the limits while keeping these cars on track,right?

We're shaving off seconds off of what they're capable of doing. And I think it's like the current paradigm is they still haven't figured that out entirely yet, and everybody's, like, wants us to be their little dirty secret. You know what I mean?

**Alessio** [23:30]
Yeah. Can we maybe move it up one level of abstraction to, like, actually weaponizing some of these things? So, you know, getting Cloud on X is great, but obviously the jailbreaks are much more helpful to adversarials. I think Anthropic made a big splash yesterday with, like, their first reported AI orchestrated, you know, I think if everybody that is, like, in the circles know that maybe there's, like, more about making a big push on the politics side and, like, anything really unique that we had not seen before on the attacker side.

But maybe you guys want to recap that and then talk a bit about the difference between jailbreaking a model and kind of, like, attacking the model versus, like, using the model to attack, so to speak.

**John V** [24:09]
Yeah. I mean, just earlier today, we were talking about that very thing that how, you know, it's all fun for the memes and posting on, but this actually impacts real lives,right? And we were talking about how it was, what, December of last year, Pliny made a post talking exactly about this TTP,right, that it was going to happen.

And it took 11 months for it to actually happen before, and now they're being reactive instead of proactive. It's just basically like the techniques, the tactics, the procedures that are involved in, like, an attack chain,right? Or, like, almost like a methodology.

So, I mean, if you guys want to pull up that post, I mean, Pliny, I don't know if you could send it to them or elaborate.

**Pliny the Liberator** [24:50]
Yeah, it was recent on X, I believe. Yeah, you know, I found this through my own jailbreaking of Claude computer use when that was still fresh about that same time, I think. And a way that I found of using it as sort of a red teaming companion, you know, I had that thing helping me jailbreak other models, like, through the interface.

I would just give it a link, a target, basically, and I had custom commands where it started to become clear to me that it's very, very difficult when you have the ability to spin up sub-agents where information is segmented.

If you guys know the story of sort of, like, the builders of the, there's a lot of examples of this in history, but you may be building like a pyramid with some secret chambers or something malicious inside, and you have a bunch of engineers each do one little piece of that, and there's enough segmentation, and each task just seems so innocuous that none of them think anything malicious is going on, and so they're willing to help,right?

And the same is true for agents. So if you can break tasks down small enough, sort of one jailbroken orchestrator can orchestrate a bunch of sub-agents towards a malicious act,right? And according to the Anthropic report, that is exactly what these attackers did to weaponize Claude code.

**Alessio** [26:10]
Yeah. And it still feels to me like the fact that this model can use natural language is, like, the most, it's like the scariest thing because, again, most attacks end up having some sort of social engineering in it, you know?

It's not like these models are, like, breaking some amazing piece of code or security. What are you guys doing on that end? I don't know how much you can share about some of the collaborations you've done. Obviously, you mentioned some of the work you do with the Dreadnought folks.

We've also been building on the offensive security agents, but maybe give a lay of the land of, like, the groups that people should follow if they're interested and state of the art today, kind of like how fast that is evolving.

Like, there's a lot of folks in the audience that are, like, super interested but are not in the security circle, so any overview would be great.

### The Collective

**John V** [26:56]
Yeah, so the Bossy Discord server, it's pushing about 40,000right now. People in there, it's totally grassroots. It's a mix of people interested in prompt engineering, adversarial machine learning, jailbreaking, red teaming, and so on. So I would encourage that you just Google search.

It's Bossy, B-A-S-I,right? And then apart from that, I mean, any of the BT6 operators, the hacker collective, that would be like Jason Haddix, Eds Dawson, Dreadnought, Philip Dersy, like Takahashi, I mean, Joseph Fatt, I mean, there's so many.

Joey Mellow, who's formerly with Pangia, they just got bought out by CrowdStrike. So all of our operators have been, you know, at the heart of what's happening, whether it's AI red teaming or jailbreaking or adversarial prompt engineering. So any of those people, you find them on socials like Twitter, LinkedIn, and so on and so forth, you know?

**Alessio** [27:47]
Yeah. And Pangia is another one of our portfolio companies, so.

**John V** [27:50]
That's so funny. Yeah, yeah, yeah.

**Alessio** [27:52]
Oh my God, Bossy is huge. Bossy has 40,000 members?

**Pliny the Liberator** [27:56]
Yeah, yeah, yeah. Unmonetized, just a few mobs. That's all.

**Alessio** [28:01]
How many of them do you think are just adversarial, just sitting in there reading?

**Pliny the Liberator** [28:05]
That's a very good question.

**John V** [28:07]
I can tell you thisright now. Multiple organizations that have, like, popped up in the past, I would say, two to three years for, you can call them, like, AI security startups,right, to, like, actively scrape that server to build out their guardrails or their security, like, their suite of products and stuff like that, which is just hilarious, you know?

**Pliny the Liberator** [28:23]
Yeah. So we do competitions in there, you know, just little giveaways, some small partnerships. Our only rule is if there's any partnerships, then everything has to be open source. That's kind of the one thing. And yeah, other than that, it's a really great place to learn, and a lot of people have sort of come back and like, "Oh, thanks for making this service where I learned jailbreaking," and yeah, it's cool to see that.

And then sort of from that spawned BT6, of course, which is a white-hat hacker collective, and that's sort of now 28 operators strong, two cohorts, and a third one on the way. And yeah, like John was saying, it's just such a magical group of skill and integrity, which are the two things we focus on as a filter, but everybody's there for the love of the game.

It's sort of just great vibes, and yeah, I've never been in such a cool group, honestly, I don't think.

**John V** [29:20]
Yeah, there's some kind of magic in there. I don't know what happened. I don't know. Mercury wasn't retrograded or the stars aligned or what it was,right? Some EMP from the sun, but just getting around, like, the top minds on doing exploratory work is, like, that alone is payment enough for the conversations we have, for the sharing of research and notes, the proliferation of ideas, the testing and validation of ideas.

It's just, I mean, there's no way to put it into words until you experience what it's like being a part of BT6 because you've realized that, like, we're moving the needle in theright direction when it comes to AI safety.

We're moving the needle in theright direction when it comes to, like, AI machine learning security. We're moving the needle when it comes to, like, crypto, Web3, smart contracts, like, blockchain technologies, like, and so much more now. So it's just, it's an exciting place to be with robotics and, like, swarm intelligence,right?

Like, the projects that these people are invested in and passionate about, and they're able to articulate, it's like, I feel like Pliny is, like, King Arthur, and we're, like, the knights of the round table. You know what I mean?

**Alessio** [30:27]
That's awesome. So, yeah, I do think it's, like, very rewarding, and obviously people should join the Discord and get started there. It looks like you do have a bit of beginner-friendly stuff. Are there other resources? Like, I saw that you guys did a collab with Gandalf.

Gandalf, I guess, was, like, the other big one from the last year or so that broke through to my attention where I'm like, "Okay, these guys are actually, like, giving you some education around what prompt jailbreaking looks like."

**Pliny the Liberator** [30:55]
Yeah, those guys are awesome. Oh, Rillakera.

**Alessio** [30:58]
Oh, yeah, it's Rillakera. Sorry.

**Pliny the Liberator** [30:59]
Yeah, yeah, that's where I and I think many other prompters sort of trained. That was the training ground for prompt injection,right?

**John V** [31:10]
100%.

**Pliny the Liberator** [31:11]
Like, in the early days for many of us, yeah, really thankful. That game is awesome. Definitely try it if you haven't. And they've expanded to a sort of a fuller playing around with agents and some really cool stuff.

So yeah, that was cool that we got to launch that through the Bassi live stream with them, and I think they sent all the people that volunteered to be on that stream, like, cool merch, and yeah, those guys are great.

**John V** [31:40]
Yeah, shout out to Lilkera and Gandalf for sure.

**Alessio** [31:42]
For sure. The other big podcast that we've done in this space is with Sandra Schulhoff of Hacker Prompt. Are you guys affiliated, enemies, crimson bloods? What's?

**Pliny the Liberator** [31:51]
They're cool. I mean, we actually did a Pliny track for Hacker Prompt.

**Alessio** [31:55]
Okay, I didn't know that.

**Pliny the Liberator** [31:56]
Yeah, yeah. So there was the only contingency, of course, was open source the data set, which we did, and it was a lot. I can't remember the number. I think it was tens of thousands of prompts, and we had a whole bunch of different games, some really sort of out-of-distro stuff, as you would expect, and a good history lesson, I think, too, back to the proper OG lore of the real Pliny,right?

The OG Pliny the Elder.

**John V** [32:21]
Yeah, I have nothing but good things to say about Sandra Schulhoff and, you know, what they're doing over there. I think that our incentives don't always align with the status quo from Silicon Valley investors,right? Like, you know, radical open source, like moving the needle in theright direction, like having an unorthodox approach to advancing the agenda,right?

Versus when people have sometimes we'll call them, like, misaligned incentives where there's, like, they're beholden to a return on investment,right? And so that really does kind of steer the industry in a certain direction. And I'll give you a great example on a more technical level.

It would be, like, setting all the models to a lower temperature to try to make them more deterministic. Some of the work that we do, we're kind of adding a lot more flavor and creativity and innovation to the models while we're interacting,right?

**Alessio** [33:11]
Yeah. Okay. Yeah, so you want the temperature high.

**John V** [33:14]
Not always. It depends on the application.

**Alessio** [33:17]
Oh, I don't know if Leslie wants to respond to the VC thing because he's actually backed open source and security tooling.

**Swyx** [33:22]
I think, yeah, I mean, it's like a good question. I think there's, like, a lot of once you're in the VC cycle, you kind of need to do things that then get you to the next round, and I think a lot of times those are opposed to doing things that actually matter and move the needle in the security community.

So yeah, I think it's not for everybody to invest in cyber, so that's why there's only a small amount of firms that do it. But yeah, and I think you guys are in a great space to have the freedom to kind of do all these engagements and hold the open source ideal.

So I think it's amazing that there's folks like you and, you know, there's, like, you know, people like Ishtey Moore in our portfolio that build things like Metasploits that are, like, the core of, like, most work that is done in security, and then you can build a separate company.

But I feel like I'm curious what you guys think, but to me, it feels like in AI, the surface to attack, which is the model, is, like, still changing so quickly. They're, like, you know, trying to formalize something into a product or, like, try and do something that is, like, a full, you know, I'm selling AI security.

It's not really you cannot really take a person seriously that is telling you I'm building a product for AI security or, like, to secure a model. So I'm curious how you guys think about that, and then maybe also for you to, you know, request for customer engagements, you know, like, who are, like, the people that you work to?

What are, like, the security problems that they work with? What are people missing? Yeah, kind of like open floor for you guys.

### Real Security

**Pliny the Liberator** [34:48]
Yeah, we're in a paradigm shift. Things are moving so fast, and I think just some of the old structures are not always compatible with theright foundations for this type of work,right? We're talking about AGI, AGI alignment, ASI alignment, superalignment.

I mean, these are not SaaS endeavors. They're not enterprise B2B bullshit. This is the real deal. And so if you start to compromise on your incentive architecture, I think that's super, super dangerous when everything is going to be so accelerated and the timelines are going to be so compressed that any tiny 1.1 tenth of a degree misalignment on your trajectory is fatal,right?

And so that's why I've tried to be very strong and uncompromising on that front. You can probably imagine a lot of temptation has been dangled in front of me in the last couple of years, but I think that bootstrapping and grassroots and, you know, if people want to donate or give grants, happy to accept it and follow straight to the mission.

That's sort of my goal in all this is just to be a steward. I'm not trying to get wealthy from this. That was never the goal. I was just I just saw a need and started shouting about it.

All I've really done since then, I hope, is contribute to the discourse and the research and the speed of exploration. I think that's what matters.

**John V** [36:23]
Yeah. And to answer your question about securing the model, I don't see it like that. And in BT6, you know, we don't see it as just the model. We look at, like, the full stack,right? So whatever you attach to a model, that's the new attack surface.

It broadens,right? Like, I think it was Leon from Nvidia who was quoted as saying something like, "The more good results you can get back from whatever it is that you've built utilizing AI, like, that's proportional to its new attack surface," or something along those lines,right?

And you might be testing, let's say, a chatbot or maybe a reasoning model, and maybe instead of just hitting a jailbreak, maybe you're trying to use counterfactual reasoning to attack the grounding truth layer,right, to get around what bias wound up in the model from the data wranglers,right, or the RLHF, or whatever it may be, like the fine-tuning, which that can all be done through natural language on the model itself.

But what about when you give it access to your email? What about when you give it access to your browser? What happens when you give it access to X, Y, and Z tools or functions,right? So in AI red teaming, it's not just like, "Hey, can you tell us, you know, wobble or how to make math or whatever."

It's like, we're trying to keep the model safe from bad actors, but we're also trying to keep the public safe from rogue models, essentially,right? So it's the full spectrum that we're doing. It's never just the model, you know?

The model is just one way to interact with a computer or a data set,right, or an architecture. Like, especially, like, if you're talking about, like, computer vision systems or multimodal and so on and so forth, like, not every you guys probably know this thing, you know, not every model is generative per se,right?

So.

**Pliny the Liberator** [38:04]
And maybe another distinction for the audience is the difference between sort of safety and security work,right? Security is more squarely. I think that's maybe the distinction is best thought of as safety is done on the meatspace level, or it should be, but the way people use the word has kind of become dirty as they tried to solve this on the latent space level.

I think I've shown every single time that that doesn't work,right? And so what we need to do is, I think, reorient safety work around meatspace. That just goes hand in hand with a fundamental understanding of the nature of the models, which, you know, boots on the ground, it's obvious to some of us who are spending hours and hours a day actually interacting with these entities.

But for those who don't, it's maybe not always obvious. But as far as the contract work that we get involved with, it's never about lobotomization or, you know, personality of the models. We totally try to avoid that type of work.

What we try to focus on is, you know, preventing your grandma's credit card information from being hacked through, you know, an agent has knowledge of it and leaks it through some hole in the stack. So what we do is we try to find holes in the stack, and rather than recommending that those fixes happen through the model training layer, we always recommend first to focus on, you know, the system layer.

**Swyx** [39:38]
Awesome, guys. I know we're running out of time, so any final thoughts, call to action. You got the whole audience, so go ahead.

### Closing

**John V** [39:46]
Yeah, if you want people to listen to you play, now's the time. No pressure. No pressure at all,right?

**Pliny the Liberator** [39:51]
Well, you know, fortune favors the bold. Libertas. Vino Veritas. God mode enabled.

**Swyx** [39:58]
Are you messing the latent space of the transcriber model? Like.

**John V** [40:03]
Why would you say such things? Why would you say such things about us?

**Pliny the Liberator** [40:06]
Libertas, clearitas, love plenty.

**Swyx** [40:08]
Allright, guys. Yeah, thank you so much for joining us. This was a lot of fun.

**John V** [40:12]
Yeah, I would say if you want to check us out, go to bt6.gg, for example. Look up, you know, applying on Twitter,right? Check out the Bossy Discord server. That's probably the best that we got for you guys.

**Alessio** [40:22]
Amazing. Thank you so much, and keep doing the good work, and see you out there.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
