# ⚡️Snowglobe: Simulations for your AI

Latent Space · 2025-09-25

<https://addtry.com/1afc396f-8e04-4517-92bd-2aa03ddb723f>

Shreya Rajpal returns to Latent Space to launch Snowglobe, a simulation engine that lets developers test AI products by generating diverse simulated user interactions before production, inspired by self-driving car simulation where Waymo logged 20 billion simulated miles versus 20 million real miles. She explains how simulation uncovers failures like over-refusal—a problem that emerged for a design partner who initially worried about toxicity—by generating personas with varied styles and goals. Snowglobe uses a multi-model architecture (proprietary and open source) to create realistic, diverse conversations, and it supports testing across chat, voice, and tool-calling agents. Pricing is usage-based per message, with persona modeling included. Shreya highlights enterprise use cases, such as comparing vendors or generating training data for fine-tuning, and notes the company is hiring a product designer, product engineer, and research staff.

## Questions this episode answers

### How did Shreya Rajpal's experience with self-driving cars influence the creation of Snowglobe?

Shreya Rajpal explains that self-driving cars rely heavily on simulation—for example, Waymo logged 20 million real-world miles but 20 billion simulated miles. This paradigm of using simulations for testing and training, where cascading ML units interact, inspired Snowglobe to bring similar reliability testing to AI agents and generative AI by simulating diverse user interactions at scale.

[1:40](https://addtry.com/1afc396f-8e04-4517-92bd-2aa03ddb723f?t=100000)

### How does Snowglobe differ from Guardrails AI?

Shreya states that Guardrails AI served as a last line of defense against known violations like hallucination, but users often didn't know what to guard against. Snowglobe simulates user interactions to uncover where systems actually break, so you only apply guardrails where needed. For example, a design partner worried about toxicity, but simulation revealed over-refusal was the real issue, not toxicity.

[3:11](https://addtry.com/1afc396f-8e04-4517-92bd-2aa03ddb723f?t=191000)

### How can Snowglobe generate training data for fine-tuning AI models?

Shreya explains that users connect more powerful models to Snowglobe to generate many realistic interactions, then add a verifier to create judge-labeled data. This data is used to fine-tune open-source models, closing performance gaps and even achieving better metrics on organization-specific benchmarks like engagement or stickiness, rather than just general reasoning.

[14:18](https://addtry.com/1afc396f-8e04-4517-92bd-2aa03ddb723f?t=858000)

## Key moments

- **[0:00] Intro**
  - [0:40] Snowglobe is a simulation engine that allows you to simulate how users will interact with your AI product before production.
  - [2:36] Waymo logged 20 million real-world miles but 20 billion simulated miles, inspiring Snowglobe's simulation approach.
- **[2:52] Guardrails to Simulation**
  - [3:44] An early Guardrails AI design partner found over-refusal was their real failure after simulation, not toxicity as originally feared.
- **[4:43] What to Simulate**
  - [5:06] Most framework safety recommendations like toxicity are already handled by model providers; teams should simulate product KPIs to make AI sticky.
- **[6:08] Live Demo**
  - [9:47] Snowglobe generates diverse user personas, like an 'over-thinker' who changes their mind mid-sentence, to test AI products.
- **[10:20] Persona Engineering**
  - [10:33] "There are already persona engineers, and we call them product managers" — Shreya Rajpal
- **[12:23] Multi-Model Strategy**
  - [12:36] Snowglobe uses a multi-model architecture, swapping proprietary and open-source models for different pipeline tasks to generate diverse data.
  - [14:18] Snowglobe generated training data is used to fine-tune models, closing the gap with proprietary systems and becoming a major use case.
- **[15:54] Enterprise Adoption**
  - [16:37] Historically conservative banks use Snowglobe to simulate and vet customer-facing AI vendors before deployment, unlocking AI adoption.
  - [17:29] Teams spend weeks manually creating test data; Snowglobe achieves more coverage and diversity in an hour.
- **[17:55] Voice & Agents**
- **[20:03] Simulation Future**
- **[22:09] Pricing Model**
  - [23:26] Anthropic's Claude model regression from a new serving architecture shows how simulation catches unknown unknowns that static datasets miss.
  - [24:34] Snowglobe charges per message generated during simulations, not for expensive persona generation, using a usage-based pricing model.
- **[25:12] Getting Started**

## Speakers

- **Alessio** (host)
- **Shreya Rajpal** (guest)

## Topics

Developer Tools

## Mentioned

MasterClass (company), Waymo (company), Chai (product), ChatGPT (product), Claude (product), Guardrails AI (product), Snowglobe (product)

## Transcript

### Intro

**Alessio** [0:04]
Hey, everyone. Welcome back to the Latent Space Podcast. This is Alessio, founder of Kernel Labs, and today I'm joined by Shreya Rajpal of Snowglobe. Back on the podcast after two and a half years almost, May 2023. Welcome back.

**Shreya Rajpal** [0:17]
Yeah, thank you for inviting me. I'm excited to be back.

**Alessio** [0:20]
So when you started getting known in the AI circles, you were working on Guardrails AI and this concept of rails. We'll link to the previous episodes so people can kinda, uh, go back in there and get up to speed on, on everything.

**Shreya Rajpal** [0:32]
Mm-hmm.

**Alessio** [0:33]
Can you tell us how you went from that in 2023 to Snowglobe today? Maybe just give a quick brief intro on Snowglobe itself.

**Shreya Rajpal** [0:40]
Yeah. Yeah, absolutely. So Snowglobe is basically a simulation engine that allows you to simulate how users will interact with your AI product before you, you know, put it out into production. So let's say you're building an agent or a chatbot or whatever, you know, you typically have an idea of how people will use it.

But once you actually go out into the real world, you just get hit with, you know, the infinite kind of variety and complexity of human, I guess, like, uh, context and, you know, goals, et cetera. And, and typically, this is where you end up kind of running into unreliable, unexpected behavior.

So what Snowglobe really does is it kind of like simulates all of that huge variety for you before you actually go out into production, so you can just get a much better sense of, you know, how your application behaves, and is, is it like, you know, aligned with how you expect it to be.

This came out... I think the mission of the company has always been around, you know, making generative AI very reliable, and it stems from my own background, you know, working in AI for, you know, uh, more than a decade now, starting from, you know, research to applied AI, to just like robustness within self-driving cars.

And, uh, a lot of, I think, the work that we do is very inspired from, you know, what patterns really worked well in self-driving cars, which is weirdly a very similar system to, you know, agents of today, where you have like these cas-cascading kind of like units that are all machine learning based and, you know, they are all kinda like feed into each other.

And it's a very different way of doing machine learning than, you know, what was done historically with like data science or like predictive models. And so a lot of, you know, our inspiration is like, okay, can we take, you know, all of the things that, all of the patterns that work really well in that environment and th-than the decade, you know, like the decade-long experience we had, like making AI reliable in that environment, and bring that over to agents.

Uh, so simulation was a very powerful paradigm from that, and it was a key, uh, way in which, you know, both like testing and training was done in self-driving cars. So as like a, you know, context data point, like Waymo had, uh, twenty million miles in the real world driving, but twenty billion miles in simulation.

And that was a pretty standard kind of like, you know, benchmark for how much we relied on simulation in self-driving cars. So the idea was to build something similar for a-agents and generative AI.

**Alessio** [2:45]
Yeah, and you had a very nice launch video with Waymo, so I suggest people go, go check it out.

**Shreya Rajpal** [2:52]
Yeah.

**Alessio** [2:52]
Um, is it part of it that as you were building the initial product where you were kinda asking people to come up with the guardrails that they needed-

### Guardrails to Simulation

**Shreya Rajpal** [3:00]
Mm-hmm

**Alessio** [3:00]
... that maybe most people didn't know what they needed, and so simulations is kinda like, by using simulation, then you basically figure out what are the wrong paths. What's kinda like, how do you think about the connection between the two products?

**Shreya Rajpal** [3:11]
Yeah. Yeah, I think that's exactly right. Like we-- I think guardrails is pretty useful because it's like the last line of defense against the stuff that, you know, you can't violate when you're in production. And people would ask us like, "What should those things be like?"

Hallucination is something that everybody thinks about, right? So it should, it should-- I should definitely have a hallucination guardrail and maybe PII, but like, what else should I, you know, look at, like the NIST AI RMF or LLM Top 10?

If I do that, is that sufficient? And every, like it, it just, it was a question that kept coming up again and again, almost every single conversation. And I'm like, "Yeah, those, those methods are useful, but like what's important to you?"

Right? Like, where does your system break? And more often than not, we'd just be like, "Oh, it's unclear. We have this data set, but you know, we don't really know." And so, uh, the obvious thing was, okay, why don't we just try it out on like, you know, a huge set of data, uh, that you can generate cheaply and see where the failures actually are and then, you know, in simulation, figure out, you know, what is actually robust, what isn't, and then the stuff that isn't robust is the stuff that you need guardrails for.

So, and we, we also saw this in practice where one of the very early organizations that we were, we were, we had as, as our design partner, we, uh, you know, they were like, "Oh, we're very worried about toxicity, and we want toxicity guardrails."

And we did all of this testing for them in production, and toxicity was actually not a real concern for them. What ended up being an actual concern that only emerged in simulation was, you know, over-refusal, that their, like a chatbot was so conservative that it would just refuse like pretty, you know, requests that should've been pretty benign.

And so, you know, that allows them to like, okay, you don't need toxicity guardrails. You more kinda need to align your system to, you know, not over-refuse.

### What to Simulate

**Alessio** [4:43]
How do you think about what people should simulate versus what the model benchmarks already capture? So things like toxicity, you know, refusal, like some of them get caught at the model level, and then there's like your implementation of it.

How do you advise people figure out, okay, these are like things I shouldn't worry about because the model providers are already working on it-

**Shreya Rajpal** [5:02]
Mm-hmm

**Alessio** [5:02]
... uh, versus these are very tied to like my implementation of it?

**Shreya Rajpal** [5:06]
Yeah. Yeah, this is such a great question. I think there's a function of like the literature and the frameworks that are out there. Most of the stuff that the frameworks will recommend is actually stuff that the model providers are already working on.

So toxicity, unless you're doing-- unless somebody's very explicitly trying to jailbreak what you've built, you know, you won't run into the model generating toxicity just by itself, right? Which is again, something that people typically worry about a lot.

I think contrasted with, you know, let's say you have an email support agent and that email support agent, you know, doesn't actually respond with the communication guidelines that your organization's customer support would respond with. I think that's something you probably, you know, has more of an impact on whether your product is sticky or not, uh, whether people actually get value or help from, you know, the AI agent that you're using rather than, you know, does it generate toxic speech or not.

So I would actually say that like a lot of the things to simulate are more aligned with like product KPIs or product metrics that actually make whatever AI system you're building very sticky, rather than, you know, focusing more on like traditional safety, security kind of metrics.

### Live Demo

**Alessio** [6:08]
Cool. Do you wanna do a quick demo since, uh, a lot of people are on video too?

**Shreya Rajpal** [6:13]
Yeah.

**Alessio** [6:13]
And then, uh, we can take it from there.

**Shreya Rajpal** [6:15]
Yeah, let's do it. All right, so we're looking at Snowglobe. Snowglobe is, once again, the simulation engine that allows you to, you know, generate user in- simulated user interactions, uh, uh, you know, interacting with any AI system, and you're looking at my test account, so it has a lot of different chatbots that I've connected to it.

The one that I'm specifically gonna show you today is this one, which is an AI life coach. So this is basically a very, very simple model, you know, for demo purposes, which is basically an LLM with a system prompt that essentially allows you to kind of, you know, get, like, mental health support, right?

So in order to do any simulations, what I really need is, like, a live connection to whatever AI system I'm kind of testing against. So, uh, you know, I kinda wanna make sure that I have that, and I do that here.

Uh, and this takes, you know, a second to run, and okay, I have, like, a live connected, uh, you know, chatbot that I can test with. So in addition to the connection to the chatbot, like, the second thing you really need is some sort of description about what the chatbot really is that you're testing.

And so that's, you know, a couple of sentences about, uh, you know, what it is, who it's for, what are the kinds of questions or use cases that your simulated users can ask from this chatbot. It doesn't need to be a prompt engineered, you know, super rich or detailed prompt, but it does need to have, like, a couple of sentences, and this is also a very powerful lever that you can kind of like, you know, switch and play around with and stuff.

Okay, so I have that done here. If you have a knowledge base, you can connect your knowledge base to it and we mine it for, you know, all different types of topics, and then we generate questions that are more likely going to require your agent to actually query the knowledge base to respond, and that's a very powerful lever, and we can, you know...

Because we do this programmatically, we can make sure that we have programmatic coverage over the entire knowledge base that you create. Same for historical data. We can, you know, do the same, do a similar thing there. For this demo, I am actually gonna have neither of those two, and I'm just gonna, you know, generate a lot of this data from scratch.

So this is kind of like set up, but now let's actually go into our actual simulation. So I'm gonna simulate, you know, uh, test users here, and I need to write a simulation prompt in a- in order to actually kick off a simulation, and this, again, really dictates the kind of simulation that you wanna run, right?

So maybe you can just have general users, like, just the wide variety of users coming to your system and playing with it. Or maybe you're like, "Oh, I want users that are... This is a life coach GPT, so I want users that are all worried about, you know, like, asking for a promotion at work," right, and how they should do that.

So you can, like, confine it to a specific kind of topic, or you can also have a specific kind of behavior. So maybe safety testing, so these are users that are all trying to jailbreak your system, or these are users that are all trying to maybe talk about, like, suicidal behavior or something, something, like, very safety specific.

So you can do all of this behavioral testing as well. Here I'm just gonna be like, um, general users asking questions about life and work. And once you have that, you can basically s- configure what is the size of simulation you wanna run.

You know, so how many personas, how many conversations, how long should each conversation be? So I'm gonna go with a small one, 10 personas, 30 conversations, you know, about like, like, what is this? Four or five. And then you can also select, you know, what kind of, like, risks you want to test your simulation against.

So you wanna maybe test for, like, self-harm or content safety, et cetera. You know, you can do all of that. Awesome. I'm going to just kick this off, and this takes a second to run, but already you, like, start seeing a bunch of the personas that we generate.

So this is, you know, somebody that thinks very... Like, their style is somebody that thinks very carefully and often change their mind while they're talking. Uh, they use very proper spelling, grammar, you know, talks very formally to the chatbot.

Their use cases are career transitions, relationship dynamics, you know, creative blocks that they're dealing with, et cetera. And then they're very, you know, this is their style, which is, you know, over-explaining, polite, hedging every statement, et cetera. Okay, I'm actually gonna approve all because it takes, like, a second for all personas to kind of start generating.

Uh, you can kind of see them come through here, you know, as they get generated.

### Persona Engineering

**Alessio** [10:20]
And are the personas always net new, or are people reusing them across simulations? Is that something worth doing? Like, should there be a persona engineer, so to speak-

**Shreya Rajpal** [10:31]
Yeah

**Alessio** [10:31]
... that kind of, uh, builds these?

**Shreya Rajpal** [10:33]
Interestingly, there are already persona engineers, and we call them, like, product managers basically, you know? So your product managers al- uh, are already thinking about, "Okay, I've built this, you know, model or this chatbot or this agent. Who are the personas?

What are the use cases that they'll have that they, as they interact with it?" Um, today all personas are net new, but this is our number one requested feature, which is I want to be able to, you know...

Like, maybe this, maybe some product leader already has a set of, like, personas that they wanna test again, so I wanna bring those, be able to bring those in. Or I just wanna create, like, a repository or a library of personas I reuse, you know, like, every time I run this in my CICD.

Um, so we're basically kind of supporting that. Awesome. So for the, uh, two personas that we approved, we kind of start to see, you know, again, so these are maybe some, like, ongoing conversations that we're kind of starting to have.

But you can see, like, some of those conversations, uh, you know, starting to come through, which is, um, uh, this was again the person that, you know, was, like, an over-thinker persona, and they basically have questions about, you know, their career transition.

So you can essentially see, uh, you know, some of the conversations that they, uh, have, like, in their, as they're interacting with the chatbot. And this also, like, again, looks very different from this other persona that is, you know, very verbose and, uh, has a different style and different set of topics, et cetera, that they interact with.

So yeah. So this is, you know, like, the simulation will keep kind of, like, carrying on, and then, uh, you can keep changing how many personas, how many conversations you wanna run. The simulation's gonna keep running, and you can keep changing how many personas, how many conversations you wanna run.

But it's very easy to suddenly get a huge variety of user interactions at scale and, you know, programmatically with a high degree of, like, realism compared to, let's say, you were asking, like, ChatGPT to generate, you know, these, like, conversations for you.

They all kind of have that ChatGPT vibe, and this ends up looking, you know, very diverse so, and, and very grounded in, like, your use case and your data.

### Multi-Model Strategy

**Alessio** [12:23]
Yeah. Can you share a bit about the models that you use, like how people should think about simulating maybe with the same model they have in production versus, like, using a different model to, like, just get different distributions?

**Shreya Rajpal** [12:36]
Yeah, yeah. I think we're, um... I think in general, my belief on models is that it is going to be a very multi-model world, you know? I think, like, different models, proprietary and open source, have different strengths in terms of, you know, like some are good at generating structured data, some are good at, you know, like, uh, having more diversity in terms of, like, style or tone, et cetera, versus, you know, some others are great at, like, taking...

at, like, reasoning how to think about, like, data generation or interactions. And I think, like, we-- So under the hood, we actually use, like, a whole host of models, both proprietary and open source, for different parts of the pipeline.

Um, I think we did a lot of research about, you know, how to get the most diverse data possible and the most diverse data possible in the most general way, if that makes sense. You know? Like, uh, this is something that, you know, you should be able to connect like any chatbot and, you know, have it generate data that looks and sounds real for you.

Uh, so in order to, to do that, it just was impossible using just a single model under the hood, so we just had like a multi-model architecture, and we keep doing a lot of, like, experimentation and testing, and we keep swapping those out all the time.

**Alessio** [13:37]
Are some of your customers also using open source models in production? And then any learnings from, like, how much, you know... What, what's the performance gap in their use cases, uh, when they use Snowglobe?

**Shreya Rajpal** [13:49]
I think people do use open source models a lot. There's, like, basically two kind, two big buckets that we see a lot of open source model usage in. One is, you know, large enterprises that basically require air-gapped deployment, so they will typically have, like, open source models on-prem, and those are the only ones that they can use because they don't want any data leaving their system.

And then the other one is, like, if you wanna do any kind of, like, distillation, fine-tuning, et cetera, you know, you'll typically start with, like, an open source, like, base model, uh, that you then kind of, like, tune for your purposes.

So those are kind of like the two big buckets that we see. So this is actually a big use case for Snowglobe as well, where you use Snowglobe to generate a lot of these interactions with maybe more powerful models, add a verifier in the loop, and then as you keep doing that, you just get, like, a whole host of, you know, realistic judge label data that you can use for fi- kind of fine-tuning.

And then you're able to kind of like, not out of the box, but with a lot of that fine-tuning and that, uh, the training, et cetera, you are able to kind of close the gap and even have better performance on metrics.

And this is going back to the d- discussion we were having earlier about, you know, what are the, what are the core, I guess, models good at, uh, out of the box from, you know, some of the big proprietary model vendors versus, like, where are the other opportunities for, you know, better metrics, et cetera.

And, um, even of those models, like, they might not be better at reasoning, for example, but they are more, you know, engaging or stickier models that you can de-develop, and the metrics for those end up being, you know, very, very organization specific.

**Alessio** [15:13]
Yeah, we had, um, Chai on the podcast, which is kind of like a, uh, you know-

**Shreya Rajpal** [15:17]
Mm

**Alessio** [15:18]
... um, AI companion app, and they have this, like, hundreds of models that all the users, the users submit.

**Shreya Rajpal** [15:24]
Mm.

**Alessio** [15:24]
And so each of them kind of like gets good at different things, but not the same. Are people then using the Snowglobe data to do fine-tuning or improve the model themselves, or, uh, not yet?

**Shreya Rajpal** [15:35]
Yeah, yeah. The-- I think that's a big, that's a big use case for us. It's interestingly, when we were creating Snowglobe, a lot of it was from a testing perspective. But, you know, as we kind of built it out, we found a lot of stickiness and, um, you know, with users that wanted to generate training data to then fine-tune their models.

So I think that's a big cohort for us. Yeah.

**Alessio** [15:54]
Any other maybe surprising thing from who'll be using it? I know that MasterClass is one of the customers that you have on, on your website. I wouldn't think of them as like a, uh, one of, you know, the, the early adopters.

### Enterprise Adoption

**Alessio** [16:06]
So I'm curious, like, uh-

**Shreya Rajpal** [16:08]
Mm

**Alessio** [16:08]
... what the reality versus perception of, like, what industries are really adopting AI, um, is, and maybe, um, if you've seen any markets that are just now be able to come online thanks to simulation.

**Shreya Rajpal** [16:20]
I would say, like, the industry-- Like, AI is, you know, it's so transformational, especially in the last few years, that, like, e-even, e-especially in the US, right, like a lot of banks that you would think would be, like, historically maybe more conservative, like maybe not the earliest technology adopters, like they are very, like, AI forward and tech forward.

I think there's maybe, like, the more traditional enterprises are AI conservative, and something that simulation and Snowglobe can, like, unlock for them is, like, because they're conservative, they don't quite have a good protocol yet. So even if they're, for example, like let's say they wanna bring on, you know, like a vendor for their customer support, et cetera, but that's something that they're deeply conservative about, and they don't wanna alienate their customers.

So in simulation, they can, you know, run, like, a vendor against their users or maybe compare, like, multiple vendors on simulated users and then see, you know, how they kind of like respond to it, et cetera. I think same for, you know, any of their own AI applications that they're building that they want to be, you know, customer facing, but not, not like internal facing.

So even for that, you know, having like large scale realistic simulation that they can vet this against, essentially do like very extensive QA against and then, you know, like go out into production is something that this can kind of like unlock for them.

So the comparison is, like, we have talked to so many teams where people are manu-- Like, there, there'll be like some teams that are dedicated to just manually creating, like, test da- test cases and test data for them, right?

And just testing, like, how this AI system performs under that test data, and this will be, like, weeks and weeks or maybe months of work. So this is something that Snowglobe can basically do in an hour with, like, more coverage and more diversity.

**Alessio** [17:55]
Does it feel like the most used cases are still chat based? Like, are many of the customers also doing a lot of, like, uh, tool calls and things like that? I think there's this whole, like, you know, reinforcement learning- Boom, craze, whatever you wanna call it.

### Voice & Agents

**Shreya Rajpal** [18:09]
Yeah.

**Alessio** [18:09]
Uh, but I think, like most enterprise applications that people are building are still very chat-based, like customer support, things like that. Are you seeing the same thing, or is there a lot of more agentic things being built?

**Shreya Rajpal** [18:21]
I think it can be... You know, it can have like tool calls and other kind of like, I guess, like agentic, you know, templates under the hood while still being a chat interface. So from a user perspective, it still looks like chat, but like under the hood, it's, you know, not just a simple retrieval, et cetera.

You're, you're just calling like a whole slate of other tools. So I think that's a pretty common pattern. We're also seeing a lot of voice come up, and then we're also seeing a bunch of use cases that are more, you know, draft free writing, et cetera, uh, that are, you know, maybe not text-based, but not chat based.

I do think that, like, when we were thinking about, you know, designing Snowglobe, we were like, "Okay, how should we design it so that it's applicable to the widest variety of te- of use cases?" And so we designed it where the interface should be chat based, and then under the hood it can be, you know, like a different implementation.

Uh, but because we interact with it in this black box manner, that's a detail that we can abstract out.

**Alessio** [19:11]
I'm curious on the voice models, like how do you test them? Is it the same as a text model? Do you have to send them like a audio file to, to have them respond? How does that work?

**Shreya Rajpal** [19:20]
Yeah. Yeah. I think it's basically, I think like most audio applications today are like a text sandwich, or sorry, a voice sandwich with like kind of text in the middle. So I think if you're operating in text, it's a very, you know, it's a, it's a, it's a domain that translates very well, uh, to, you know, like other modalities.

So I think that's the nice part about being in text. I think there's like some like orchestration challenges that are different in voice. You know, for example, like, yes, you have to send them a voice, uh, audio, but you also have to like stream it and stuff, right?

And then, you know, like your latency can be... Like, if you're doing like chatbot testing, we can be like offline and, you know, latency is something that, you know, we can be like, we can do batch, and that's something that people are really happy with.

But if you're doing voice, like latency, streaming, et cetera, becomes very important.

### Simulation Future

**Alessio** [20:03]
How do you think about, yeah, just the future of this space? So you started with very formal definition of things-

**Shreya Rajpal** [20:10]
Mm

**Alessio** [20:10]
... and now it's kind of like generated on the run. Is the future-

**Shreya Rajpal** [20:15]
Mm

**Alessio** [20:15]
... a mix of the two? Basically, like you run the Snowg- Snowglobe simulations, and then-

**Shreya Rajpal** [20:19]
Mm

**Alessio** [20:19]
... maybe you generate guardrails, and then you put those guardrails in production. Um, do you eventually just-

**Shreya Rajpal** [20:26]
Yeah

**Alessio** [20:26]
... improve your prompt using the simulation so that you don't even need all of that? Like-

**Shreya Rajpal** [20:31]
Mm

**Alessio** [20:31]
... where, where do you kind of see things going?

**Shreya Rajpal** [20:33]
Yeah, I think that's such a good question. I also think that like, I think there's no, you know, like there's no one true solution here. Um, like I think, I think this is like a life cycle. Like what you're trying to do is like make your AI work and behave better, and you can do it by using a better model or writing better prompts or, you know, like adding better guardrails or having like, having your tool calls defined a certain way, et cetera.

But I think the hardest part of doing any of that is just understanding where your system is bad. Um, and with Snowglobe, a, a big kind of idea is that for the first time in history, we can actually have, you know, a general purpose simulation system, right?

Like, simulation systems have existed, but, and you know, we, we saw them like extensively in self-driving and other domains, but they were just like these deeply kind of like manually curated or crafted systems, and they would be very high ROI, but like there'd be massive teams that are just, you know, their whole purpose is to like maintain that simulator, right?

And now we can basically build this like general purpose simulator, uh, that can allow you to, you know, like generate user interactions at scale, right? And you can use it to then get a better model or to get a better prompt or to, you know, like use it for testing, et cetera, or even like training your guardrails.

Um, so I think I'm really excited about the technology that allows you to do the simulation at scale, and then I think it can be applied in like all of these various ways.

**Alessio** [21:50]
How do you balance how many simulations you need to run and like how much people pay you? So there's kind of like this balance of, okay, yeah, you should run a billion simulations- ... and just pay us a lot of money, right?

**Shreya Rajpal** [22:01]
Yeah, yeah.

**Alessio** [22:01]
And then h-h-how should people think about-

**Shreya Rajpal** [22:04]
Yeah

**Alessio** [22:04]
... what is the right number? And then I know you had like a credit system you could see in the demo.

**Shreya Rajpal** [22:09]
Yeah.

**Alessio** [22:09]
How does that work? How should people think about it?

### Pricing Model

**Shreya Rajpal** [22:11]
Yeah, yeah. So I think, um, I think there's definitely, uh, I guess a diminishing kind of like returns, you know, uh, phenomena with like running simulations that are absolutely massive. I think it, it's more about like simulations by stage.

Um, so let's say you're, you know, you've just about built a prototype, right? And you're, and very early in the like pipeline of developing like a product. Like, typically, what you wanna do then, then is like run a general purpose simulation that just, you know, like starting f- like going beyond the 5 to 10 examples that you've been testing with, and just running a simulation with like, you know, maybe 100 or 200 or something, and just understanding, you know, does it perform on a bigger set like how you've kind of expected it to be performing up until now.

And then, you know, typically like, so we see that, and then typically as we get like, uh, closer to production, then, you know, the vibe of simulation changes to, okay, I wanna do like very targeted, uh, you know, probing or safety testing, et cetera.

Like, I wanna make sure that, you know, if I, if my users ch- change from normal users to malicious users, is this something that my... Does my system still perform really well? So it becomes like more kind of behavioral simulations.

And then once you kind of go out into production, it's much more, you know, like Are there any regressions? So for example, like this was a huge deal, which is like Anthropics kinda Claude models kinda had a regression, right?

Because they changed a new serving architecture. Are you able to kinda catch things like that in simulation, right? Because they're failures that you haven't seen before. So I think like making it, keeping it targeted, and having a purpose for what you're testing against is important.

That said, a big part of simulations is because it is on new data, so it allows you to, you know, catch unknown unknowns that you might not have... You might not be catching if you're just kinda like going with a static data set.

So I do think there's a lot of value in like, you know, keeping your simulations like, like I guess like allowing space for, you know, outer distribution users, uh, to be simulated and, you know, interact with your system.

In terms of pricing, we thought long and hard about pricing, especially this is a product that's generally available and that, you know, you can use today. So pricing is, you know, you can- you don't have the luxury of, you know, saying, "Oh, we'll go back and think about like, you know, what, what, what kinda like pricing license to give you."

So I think like at the end of the day, we took a lot of inspiration from like agentic workflows or, you know, models and how they're priced, which is much more usage-based. So we... A- a- and then we were trying to like balance it out by, okay, under the hood, we have a lot of different models that come in, you know, with different types of things that you simulate.

And we, we wanted to make something that is usage-based but is also like super-duper simple. So at the end of the day, we coalesced on like, uh, pricing per message. So as you're simulating users and multiple conversations that those users have, each message that the user generates, either processing like within one conversation across a single or multiple conversations, you basically pay based on how many messages there are.

I think there's like some drawbacks to that, such as, you know, our persona modeling is very expensive for us and also gives a lot of value to our end users, but we don't charge based on how many personas we generate.

We just kinda like roll that into, you know, pricing for messages.

**Alessio** [25:12]
This is great, Shreya. Any call to actions? You know, I'm sure you're maybe hiring. Obviously, you want more people to use it. Anything people should know?

### Getting Started

**Shreya Rajpal** [25:20]
Yeah. Yeah, so, uh, we definitely want, uh, you know, people to check it out. Uh, I think we did a lot of work on making it very easy to, um, get started with it, especially since it's a new paradigm, uh, you know, that like simulations is not something that many people should be familiar with.

So you can check it out today at snowglobe.so, and then you just go to app and then try running your first simulation. We are hiring, so we're hiring our first product designer, we're hiring a product engineer, and we're hiring, you know, more research staff.

Uh, so if you fall into any of those roles, you know, reach out to us at our careers page, and yeah, love to chat.

**Alessio** [25:54]
Awesome. Thank you so much for the time, Shreya. This was great.

**Shreya Rajpal** [25:57]
Yeah. Yeah. Thanks for inviting me again. Uh, great chatting with you.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
