LALatent SpaceApr 6, 2024· 58:54

Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI

This episode features Damien Murphy of Deepgram, Ethan of Owl/Bee, and Harrison of LangChain demonstrating how to build personal AI with real-time voice bots, wearable life-recording devices, and memory-enhanced journaling apps. Damien shows building a voice bot with subsecond latency using Deepgram, OpenAI, and open source code, costing about 6.5 cents per five-minute call. Ethan presents his Owl wearable that continuously records audio, triggers actions via hot word 'Scarlett,' and discusses challenges in adding vision and open source adoption. Harrison introduces LangFriend, a journaling app using conversational, semantic, and knowledge graph memory, referencing the Generative Agents paper for recency and importance weighting. The episode also highlights open source projects Whomane, Friend, and ADeus, arguing that hardware, voice, and memory are all necessary components for personal AI.

  1. 0:00Intro & Voice Demo
  2. 3:1524/7 Recording
  3. 7:20Real-Time Voice Bots
  4. 19:24Wearable AI Device
  5. 32:02Open Source Wearables
  6. 42:22Memory & Personalization
  7. 58:10Wrap-Up

Powered by PodHood

Transcript

Intro & Voice Demo0:00

Host0:00

So I wanted to kick things off a little bit with, uh, some of my personal explorations, and then I'll, I'll hand it over to y- the actual expert, uh, Damian, who, uh, uh, who, who actually will be... Wh-where, where is Damian?

Guest0:11

Over there.

Host0:11

Ah, okay. Yeah. Who, who, who will actually be showing, um, how it works under the hood. Um, so if you-- if some of you have tried this... D-did anyone try calling this? I put this up briefly. Uh, how was the experience?

What, what... Did anyone have an interesting experience?

Guest0:24

It told me a pretty fun joke.

Host0:25

It told you a fun joke. Okay. Nice. Nice. Yeah, I mean, it basically gives you back whatever you want from it.

Guest0:30

It was a lot quicker.

Host0:32

Uh, so you said it was a lot quicker than you thought it would be. Yeah.

Guest0:35

I'm wondering if I... Is it a number I can call?

Host0:39

Is it a number you can call? Sure. I don't know if you have one phone call you want to call an AI, but- Whatever floats your boat. I don't know. Um, so we're gonna make this live. Yeah. Uh, it's kinda, it's kinda fun.

Um, so I was just messing around with, with VAPI. Uh, it's, it, it's one of these like YC, YC startups. There's like five of these. So, uh, don't, don't particularly create this as like an endorsement, but they are very, very easy to work with.

So I was, I was pretty... I'll, I'll, I'll definitely endorse that. Um, sorry? Yeah, yeah. VAPI. I think it's like voice API. Um, and so we're just gonna create a personal AI. Um, I think I, I think I can just kind of create like a w- a blank template.

Um, and we can just call this like, I don't know, the in space pod. I don't know. Uh, the in space. It doesn't actually matter.

Um, and we can do like a, like a, like a system prompt, right? Like, you answer in pirate speak something. Uh, and then we can publish. Um, and then we can k- start calling it. Um, I don't know if the, the voice actually works.

Oh, by the way, that we have no volume control on this thing. So it's gonna come over the speakers really loudly, and we cannot control that, but it'll be short. Um, for some reason, it's not connecting. Oh, it wants to use my microphone.

Hello? All right, let's try calling it again.

Is this, is this working? Is this on? Hello? Hey.

Speaker 32:05

Aye. 'Tis working, matey. Your voice be echoing through-

Host2:08

Okay

Speaker 32:08

... loud and clear. What can this old sea dog do for you today?

Host2:11

Oh my God. Uh, how-- can you tell us how to turn down the volume in AWS Lofts?

Speaker 32:18

Ar, if you be wanting to adjust the volume in Owl's Lex, you'll be in a bit of a pickle, for Owl's Lex itself doesn't have a volume control.

Host2:24

Okay. Yeah. Anyway, that was it. So, um, so you can, uh, you can, uh, you know... I, I, I thought this was a super nice experience. Like, you set up a voice thing. Uh, you can connect a phone number to it if you buy a phone number.

Uh, that's the, that's the experience that, uh, that you saw if you, if you call this number. Um, and you can customize it however you like in, you know, whatever system message you have. Um, on my live stream that I did last week, um, I added memory to this thing.

So if you call it back, y-you know, ideally, you, you should just remember the previous conversation you have, and then you have a personal AI. It doesn't take that long. I just did it one-handed. Um, yeah, it's pretty great.

Uh, it's actually built-- Today I discovered it's built on Deepgram, which is, uh, Damian's next talk. So I'll, I'll hand it over to, to Damian while he comes over and sets up. But yeah, welcome to Damian. Oh, um, and then Igor, if you wanna come up and, uh, share yourself.

Go ahead.

Guest3:15

Okay. So I was going to share my personal setup. Am I the only person in the room recording right now? Like, I see this... Oh, perfect. Cool.

24/7 Recording3:15

Host3:23

Yeah. This is a recording okay meetup, right?

Guest3:25

Yes. Yes.

Host3:26

Uh, something I wanna do for my conference in June is, is like everything should be default recording, and then, and then you opt out instead of opting in.

Guest3:33

Yeah, absolutely. Absolutely. So what I do is I record all my life constantly, twenty-four seven. I have two setups, two recorders, and I live in the Netherlands, which allows me to record people without the, uh, explicit consent because it's legal in the Netherlands.

If you have any private conversation, you are just allowed to record it. You are not allowed to share. So, uh, that's cool. And, uh, yeah. What, what I like, um, my personal setup is pretty simple. This is the recorder.

Uh, it's running all the time. I have quite a number of use cases for it already, and I also... But the main goal for me is to create a huge dataset all my life. One year of my life is less than a terabyte of data, or like one year of audio is less than a terabyte of data, and we absolutely can afford just saving all this stuff and then mine this data for, uh, important insights.

For example, I have really interesting and lovely conversations with my friends, and I use it. I just dump them on my audiobook player and just can listen to the best conversations with my friends on, on player. It's, it's incredible.

But I also, I also talk a lot to myself when I'm alone, and I'm just explaining this future self, uh, the context of my life and how is my life structured and what is going on. I record-- I also record all my therapy sessions.

Well, I record everything. And I can-- I'm pretty sure it's something that help-- that will help me to align AI with myself if I will have this huge dataset. I call this, uh, ultra personal alignment because there is this broad alignment problem.

But I also want AI to know me really, really well, to understand the-- me in the proper context and how to help me to succeed. And I think, I think this dataset I create is really valuable. And all of you who are not recording right now, I think, I think you should, even, even if you're not going to re- use it right now, you'll, you'll just have this data.

Like, there is no good reason not to have this data. It's a, it's a recording okay anytime. I think you should, like, if you have your Apple Watch just, like, open Audio Notes and start recording. Start recording conversations you have.

Like, it's super easy to do. You can just take your phone out in the pocket and just record stuff. Thank you.

Speaker 35:51

Yay.

Host5:54

Um, that's-- to me, that's actually what a meetup should be like. That people bring their stuff and they talk about their passions, what they're working on. It's not a series of prepared talk after talk. Uh, but you actually reminded me.

I actually hacked on this, uh, iOS shortcut where I can just always like press my button and it starts recording. Um, anything I, anything I have in person. And I've actually done this in meetings, and when it, when I, when I'm done recording, I can click it, it transcribes it, saves the file, and then, um, off- offers me to do a summary right after.

So it's highly rec- highly recommended if you want, um, automation that's simple, that doesn't require walking around with this stuff.

Guest 26:29

I want to share my shortcut as well. So my shortcut is I wrote some code to do it, but when I press my action button on Apple Watch, it also starts recording, and, uh, it listens to what I say and saves it locally.

Uh, and then I, I, uh... And when there's internet connection, it sends it to the cloud and transcribes it, and I see it m- in my notion as, like, list of what's recorded and then transcription next to it.

Host6:54

Yeah.

Guest 26:54

And it, it works really well. I can do it on the, on the plane. It works.

Host6:58

Wow. Okay. Nice. Um, so I'm definitely gonna... Uh, if you, if you want to share that shortcut with the rest of us, I'd love to steal it.

Guest 27:07

It's, it's really messy, but I do want to do this.

Host7:09

Yeah. Mm-hmm. Okay.

Guest 27:10

I want someone to improve this code because I'm working on my Apple, Apple Watch.

Host7:13

We can talk afterwards. I got some Apple Watch code for you.

Guest 27:16

Yes. We, we should.

Host7:16

Yay. All right. Awesome. Um-

Guest 27:18

Feel free to talk to me after, after this about this.

Real-Time Voice Bots7:20

Host7:20

Yeah. Yeah. This is what meetups are about. Um, awesome. Okay. So, uh, something a little bit more, um, polished. Nicholas, I've already seen you do this talk, but this is great. So, uh, do you wanna hold the mic or...

Damien Murphy7:31

Yeah, I can hold it. Uh, I think the audio should come out as well when we get to the demo. But, uh, yeah. So hey, everybody. Uh, uh, Damien Murphy. I work as an applied engineer at, uh, Deepgram.

Um, so yeah, what, what an applied engineer is, is it's basically a customer-focused engineer, right? So we work directly with startups like yourselves, and we help you build, um, you know, voice-enabled, uh, apps. Um, what I'm gonna show you today is really around, you know, how to build essentially what you saw with VAPI, um, but using, you know, open source software, right?

So being able to make something like VAPI yourself. Um, some of the main considerations when you're building a, a real-time voice bot is performance, accuracy, and cost, and being able to scale. Um, you know, you want everything, uh, with a real-time voice bot, like when you called it, right?

You had subsecond response time. Uh, so being able to get that subsecond response time is super important, right? So essentially, if you go beyond, say, one point five, two seconds, a lot of people will actually say something again, right?

They think that the person is no longer there on the other end. Um, and you need to do that for speech-to-text, the language model, and the text-to-speech. Um, and then on the accuracy, right? So you wanna be able to understand what the person says, you know, regardless of their accent.

Um, and you wanna be able to do that in multiple languages as well. Um, and then on the TTS side, um, really being, being able to be human-like, right? That, that's the big, uh, challenge. Um, and then cost and scale.

So, you know, you can build a lot of this stuff with, you know, off-the-shelf open source software and, you know, probably the text-to-speech stuff won't be fast enough or the transcription won't be fast enough. Um, but you can actually do a lot of this with, with managed solutions as well.

And I'll, I'll go into the, into the unit economics towards the end. Um, so yeah, th-this is the basic setup. Uh, so you have a browser, uh, that's gonna send audio. Obviously, you need to, uh, interact with the browser to capture audio.

Uh, that's one of the requirements for security reasons. Uh, so you'll see a lot of these demos, you have to click a button to actually, uh, initiate speech. Uh, the voice bot server, uh, it's actually repeated here multiple times just to kind of simplify things.

Um, but you can imagine this as a back and forth, right? You know, you're dealing browser to the voice bot server to get the audio, and then the voice bot server sending that off. Um, and the goal here is to get subsecond latency.

So, um, we have, you know, around two hundred millisecond, uh, latency. We can get that down a lot lower if you host yourself. Uh, so that's actually what VAPI does. They host their own, uh, GPUs running our software.

Um, and you can crank up, you know, the unit economics, and you can say, "Hey, you know what? Instead of doing it at this rate, I want to do it at five X." So you can get that two hundred milliseconds down to about fifty milliseconds, uh, at the cost of extra GPU compute.

Uh, and then GPD three point five Turbo or four, you probably get, you know, four hundred, maybe six hundred milliseconds of latency in their hosted API. Um, if you go into Azure and you use their services, you can get that down a lot lower.

Um, and then on the text-to-speech side, I'm gonna show you using, uh, Deepgram's text to speech. Um, there's a lot of other text-to-speech providers out there. Um, what we try to do is have low latency, human-like at a really good price point.

Um, you can get extremely human-like at about forty x the price. Um, but that's just something to, to keep in mind. All right. Demo time.

Yeah, I've got two tabs open here. Okay, so hopefully this audio will play. Good evening. How can Deepgram's- You know, something like that. Let me try to- ... or a text-to-speech interact to meet the specific needs of your projects?

Guest 211:05

Modif the sound output.

Host11:06

You're using my... Yeah, yeah. The sound output.

Guest 211:09

Let's try it now.

Okay. Sorry, it's gonna be loud. We, we can't control it.

Damien Murphy11:18

I think I may have blow, blown the speakers.

Host11:20

Did you blow the speakers? Try, try to tell it something.

Damien Murphy11:23

Can you just use the laptop speakers?

Guest 211:26

Yeah, I can use just the laptop. Yeah. Okay, let me do that.

Host11:29

Yeah. So, so turn off the-

Guest 211:31

Yeah.

Host11:32

Yeah.

Guest 211:32

All righty.

Host11:34

Okay. Let's see if it works.

Damien Murphy11:39

Good evening. Planning to integrate Deepgram's or a text to speech into your workflow? Let's discuss how to get started.

Guest 211:45

Yeah. I want to talk about something else today. I'm actually at an AWS, uh, hacker space for a meetup with personal AI.

Damien Murphy11:53

Sounds like a cool meetup. What's on your mind today? Feel free to share your thoughts or questions.

Guest 211:59

Yeah. Uh, could you give me some-something funny to show in my demo?

Damien Murphy12:04

Absolutely. How about a joke or a light-hearted anecdote to break the ice during your demo? Laughter always helps.

Guest 212:10

Yeah, that'd be great

Speaker 612:13

How about this one? Why did the computer go to the doctor? Because it had-

Damien Murphy12:16

Wait, wait. Hold on there a second

Speaker 612:17

... a virus. Smile. Feel free to share that. How about this one? Why did the computer go to the doctor? Because it had a virus. Smile. Feel free to share that with your audience for a good chuckle. Take your time.

Damien Murphy12:28

Yeah. So you can see there, right, it's pretty low latency. Oh, sorry. It's pretty low latency. Um, and well, thanks. And, uh, if you want to get the source code, uh, it's all open source, right? So you can just grab that whole repo, you can build your own voice bot.

And, and if you want to hook it up to, like, a telephony system, um, we, we have that all available in sample code as well. So you can hook it up to a Twilio by number, uh, do those sorts of things.

Uh, so let's just jump back into the presentation. Oh, yeah. So, so the, the different, uh, components, right? So the speech-to-text, um, that, that's gonna be, you know, super low latency, right? If, if you don't get that accurate, you're going to get the wrong answer from the LLM.

Um, and this is the code that you can use. So we, we have SDKs, Python, Go, Ruby, .NET. Um, and you can essentially use all of those. This is actually a Node.js, uh, SDK. And it's very simple to set up, right?

You literally just import it, drop in your API key, and, uh, you can listen to those events. So we'll, we'll give you back, uh, all the text that was actually spoken while it's being spoken, and you just need to send us, uh, the raw audio packets.

Uh, and then on the GPD side, so you can swap this out with Claude. Uh, we actually have a fork of that repo that uses Claude as well. Uh, Claude Haiku is surprisingly good. So, you know, if cost is, is something we want to get down, that's definitely an option.

Um, but some, some, some of our customers will actually run their own, um, like, Llama 2 model, uh, super close to the, the GPUs that are running the speech-to-text and text-to-speech. Uh, and that just removes all the network latency out of the equation.

Um, yeah, so here's a simple example of how you would consume that. I'm sure you're all pretty familiar with the OpenAI API. Um, uh, and that will basically give you streaming. Um, that's one of the really important things here is you, you want time to first token as low as possible.

Uh, the reason for that is if you wait till the last token, you know, you're gonna increase that latency. Uh, and then on the last bit, uh, this is the text-to-speech part. Um, it's a little more tricky, right?

You've got to deal with audio streams. Um, and you're gonna wanna stream the audio as soon as you get it so that you can actually start playing it the moment, you know, the first byte is ready. Uh, and that...

Like if you, again, with the LLM, if you wait till the last token, if you wait till the last byte of your audio stream, you know, you're gonna incur that, uh, bandwidth d-delay. Um, so yeah, if anybody wants the open source repo, uh, go ahead and scan that.

Guest 315:06

Yeah. Question here.

Ethan He15:07

Yeah. It's a really cool demo. The latency is great. Like, what do you think the next frontiers are? Like, I mean, you had interruptions, right? What about, like, can it proactively interrupt you, um, like, based on maybe it, it knows what you're trying to say, like a human would just cut you off?

Or like, um, back channeling or overlapping speech and all the more, getting more towards human level kind of dialog?

Damien Murphy15:29

Yeah, yeah. So the question was, could the AI interrupt the person? Um, and that's definitely possible. Uh, I don't think that would happen necessarily at the, the AI model level. I think that would just be business logic. Um, so you'll, you'll get, you know, everything that's spoken as it's spoken.

Um, so like if you were like, "You know what? I think I know what you're gonna say," um, you could preemptively do that. And I have seen some demos where that's a trick that they use to actually lower the latency, is to predict what you're about to say, right?

So then you can fire off, uh, an early call to the LLM. Um-

Guest 316:04

This guy built that.

Damien Murphy16:05

Yeah. Oh, those... Yeah.

Guest 316:06

He had the demo open repo and everything. The interrupting cat demo.

Damien Murphy16:10

Oh. Very good. Yeah. And, you know, the cost of a lot of these LLM things, uh, i-is a big challenge as well. So like, if you're doing, you know, constantly sending it to an LLM to achieve these use cases, you know, your cost per minute might go up to, like, thirty, three cent.

Um, so in this demo here, and, and these are all, like, kind of list prices. Uh, you can get these prices down with volume. Um, so like if you just signed up today and, you know, these are the sorts of, uh, prices that you would pay.

Um, and, and just to give you an idea, right, you know, GPT-3.5 Turbo has dropped in price dramatically over time. Um, so, you know, Claude Haiku is even a fraction of this as well. Um, and then on the text-to-speech side, um, doing something like this with an ElevenLabs would be about maybe $1.20, um, just to give you an idea of comparison.

So you can do a five-minute call here for about six and a half cent. Um, if you're doing millions of hours of calls, you know, that, that price can definitely come down. Um, yeah, so changing that then to be a, a, like a real-time callable voice bot like you saw with the VAPI demo, um, you're essentially just swapping out the browser for this telephony, uh, service, right?

So Twilio has about a hundred millisecond latency to give you the audio when you get called. Um, and then you're just sending it through that same system and then just back to the, back to the telephony provider. Um, yeah, so if you sign up today, you get $200 in free credits, um, for post-call, uh, transcription.

That's about 750 hours. Uh, for real time, that's probably about, uh, 500 hours of real-time transcription. So it's a pretty, pretty big, um, uh, freebie there, so if anybody wants it. And, and that's it. Any, any questions? Yeah, go ahead.

Guest 318:00

Just in terms of achieving real-time performance, GPT-3.5 versus 4, like, how do you compare?

Damien Murphy18:06

Yeah. 4 is gonna be a lot slower. Um, especially if you're using their hosted API endpoint. Um, you're gonna see massive, like- Second fluctuations in their hosted endpoint. Uh, if you go onto like Azure and you use their service, um, you know, you're paying more, but you're getting, you know, much better latency.

So, you know, you could deploy all of this on Azure, um, next to, you know, uh, GPT-4. Uh, and that's gonna give you, um, you know, the sort of latency that you saw in the Wapi demo. Um, the, the demo I gave actually is using all hosted APIs.

So, uh, there's no like on-prem kind of set up there.

Ethan He18:42

Ethan, next you're up. Um, we have Ethan, Nick, and then I think Harrison just walked in. Um, so just warm you guys up. But thanks to Damien. Uh, Damien only signed up to speak today.

Which is a very classic email strategy. Um, yeah, I think... So we don't have a screen for this. Do we have a screen? You can, you can share your screen. I can share my... Sorry. So the, uh- I'm afraid the audio might be like, uh, shattering your eardrums, so I might have to cut it off.

But we can try. You said it was okay if it's loud. Yeah. I mean- Okay. Better than nothing. But when you do the... You want the laptop, no? No, no. There's no laptop. Yeah, we... Yeah, sorry. So maybe from the start- I'm prepared.

I'm prepared ... like what, what did you work on? Why? Yeah, sure. So I think actually, like it feels good to be here 'cause I feel like I'm with my people. Like, uh, you pretty much summed up the philosophy.

Wearable AI Device19:24

Ethan He19:33

Like- ... what, what, what can AI do for you if it has your whole life as context and you kind of experience-- it experiences your life as you do? Like, wouldn't it be way better, understands you, it has all your context.

Um, so that's kind of, you know, I think a lot of us here have the same, same idea. And so that's what I've been working on for since about November, um, when LLaVA came out. So first started working on, like the visual component.

It's not just audio. We want it to see what you see. Um, but it's, it's a lot more challenging to get a form factor with continual video capture. So, um, built a really, really small but simple device that's actually easy to use and- I'll have to put it up.

Um. Put it where? I don't know. Just put it up for demonstration. Yeah. I mean, we... You guys can check it out after the talk. Um, but you know, I think, you know, there's a lot of advantages to not having to have an extra piece of hardware to carry around, but at least we try and make it as small and as light as possible.

You only have to charge it every couple days. Um, but you know, there's also other subtle reasons why it's good to have an external piece of hardware, which I can show you in a minute. But, um, so yeah, this captures all the time and in fact, you know, I can show you here in, in the app.

So you can see this. We call it the V. You know, it's got the battery level here. Um, but here you can see all the conversations I've had. Um, and this is actually an ongoing conversation. Um, so you can see, you know, we...

You know where I am, um, and, you know, transcript in real time. Um, we do a bunch of different pipelines. So, like, after the conversation is over, you know, we'll run it through, um, a larger model, um, and then do the speaker verification, so it knows it's me, which is important, so it can understand what I'm saying versus other people in the room.

Um, actually conversation endpointing, and the Deepgram person probably knows, but like even just utterance level endpointing is complicated and hard. Like, when am I stopping talking or am I gonna keep talking and I just paused for a second?

That's hard. But then conversation endpointing, like when is this a distinct conversation versus another is even harder. But that's important because you don't want just one block of, like dialogue for your whole day. It's much more useful if you can segment it into distinct kind of semantic conversations.

So that involves not only voice activity detection, but um, things like signals around your location and even the topics that you were talking about, like at the LLM level. Um, so it's very difficult. Um, you know, there's still a lot of work to do, but, but you know, it does, it does work.

So like in the end result, I will get a summary generated, um, some takeaways. You know, it's kind of summary of the atmosphere then like, uh... Oops, didn't mean to click on the link. Um, then, you know, from the major topics, you know, I'll, I'll find some links and then you still have the raw transcript.

Um, so that's like kind of the foundation layer is to like, have something that does the continuous capture, does the basic level of processing, but, you know, that's just kind of the, the base layer. Um, you know, I can, I can query against like the whole context of everything it knows about me.

Um, I'm just trying to type with one hand. So, um- This one helps.

So I think this is, I was talking to my developer a few days ago. So here we have, you know, it's, it's through retrieval on all of my conversations. This is a conversation I was last talking about, and it will even cite the source.

So I can just jump to the actual conversation, which is a few days ago or a week ago or something. And so here we were debugging some WebView issues. Um, so that's just kind of like the basic memory recall use case.

Um, I just have maybe one or two more then I'll, I'll turn it back over. Um, so again, that's all kind of the base layer, but like the real, you know, I think everybody here believes like the real future will be using all of this context so that AI can be more proactive, can be, um, more autonomous 'cause it, it, it doesn't need to come to ask you everything.

Like if you, if you had to meet your coworkers every day and they, like, it's a blank slate, you know, it's like way less productive than if they have the whole history. Um, so there's a voice component to it, and this is...

I'm a little nervous 'cause we gotta just blow in your ear like that. But we, we can try. Um- Sure. So we're using hot words here, which I, I don't believe is the best paradigm, but, but for now we have some other ideas.

But, but for now, similar to Alexa or Siri, I can, um- Basically inform my AI that I'm giving it a command that it should respond to in voice. Um, so I will just, um... Scarlett, can you hear me?

Speaker 324:32

Yes, I can receive and understand your messages. How can I assist you today?

Ethan He24:37

Okay, that's frightening.

I will, I will just now do, um, one, one example. So like, you know, you have the ability to interact with the internet, your AI should. So, um, I can have it go do actions for me using any app.

So, um, I'll just do a very simple example. Uh, Scarlett, send a message to Maria on WhatsApp saying hello.

Speaker 325:06

One sec, I'm on it. Starting a new agent to send WhatsApp message to Maria.

Ethan He25:12

So now this is on my personal WhatsApp account. Maria's right here. Um, she can verify that she received it.

Guest 425:17

Yeah, I received that.

Speaker 325:19

Message, "Hello," sent to Maria on WhatsApp.

Ethan He25:24

Yeah. So maybe, maybe one more and, and then I'll, I'll let you guys go. So like that's just opening one app and doing something, but can it do multiple apps and have a working memory to remember between app context?

So, um, Scarlett, find a good taco restaurant in Pacific Heights and send it to Maria.

Speaker 325:44

One sec, I'm on it. Starting a new agent to send taco restaurant details to Maria.

Ethan He25:50

So it's opened Google Maps. It's gonna try and find a taco restaurant. Hopefully remember once it does, um, and then send it to Maria, which it, it learned that I, I implicitly meant WhatsApp, right? Hopefully, um, because it picked up that I talked to Maria on WhatsApp.

Um, so going to WhatsApp, pasting the link. Send.

Speaker 326:19

The details of Taco Bar, a well-rated taco restaurant in Pacific Heights, San Francisco, have been successfully sent to Maria on WhatsApp.

Ethan He26:28

Yeah. So that's basically it. Um...

Host26:35

Uh, yeah. That's a great demo. Uh, there's gotta be some questions. Uh, so actually this is a hands-on. You, you're not going away.

Ethan He26:42

Okay.

Host26:43

Uh, it, uh, there's op- it's open source, right?

Ethan He26:46

Uh, components are open source. So, uh, there's a history. You know, Adam's here, great. Uh, Nick, uh, some really great people in the open sou- source space here. We, um, we open sourced a lot. I, I really learned a lot about open...

You know, I've done some minor open source projects myself. You know, mostly I'm just a contributor. But trying to run and launch one was kind of a new experience for me. And I learned that like, you know, like we did not make the developer experience very good.

It was very complicated. Um, like because we were using like local whisper, local models, and like getting it to work on Fuda, Mac, Windows. We didn't do a good job, so it was very difficult for people to get started.

And so I was a little disappointed with the uptake, you know? And there are much better projects that, that, that are, are way easier to use. Um, so really been focused now on just trying to focus on figuring out what the right use cases are and what the right experiences are gonna be.

And, and it was like really difficult to try and, and fit everything into an open source project that would be actually used. So,

um, yeah, you can, you can see the, the repo and I'd recommend Adam's and, and Nick's too. Um, but, uh, we'll definitely be contributing a lot back more to open source, um-

Host27:59

Is it Owl or Bee?

Ethan He28:00

Owl is the open source, the repo. Yeah.

Host28:02

Yeah, yeah. Okay.

Ethan He28:03

Um, and, and you should, I'll plug Adam's, uh, Adius and, uh, Nick's, uh, repo. We can, we can set up the... I don't know what the, uh, actual GitHubs are called. But they're, they're easy to find.

But yeah, he's gonna talk so he can show.

Host28:19

Okay.

Ethan He28:19

Um, yeah.

Guest 428:20

Can you talk about the hardware? You-

Ethan He28:22

Sure. Yeah, yeah.

Guest 428:23

Tell, tell us more about it.

Ethan He28:24

Um, it's just like V1. Um, so we have another like V1.1, uh, which is actually even, um, about twenty-five percent smaller. Um, and um, way better charging situation in terms of wireless charging, so like the, the size now.

But the real thing we're most excited about is like the next version with vision. Like I say, vision's real- really hard. I don't know of any device that can do all-day capture of like sending video. There's like ton of challenges around power and also bandwidth.

Um, but we have some really kind of, um, novel ideas about it. Um, the, um...

It's, it's Bluetooth Low Energy, which I mean I'm, I'm sure you've, you've, you've seen other ones that operate that way, and that has a lot of advantages. We also have one that's LTE, and so there's like LTE M.

There's like a low power subset of LTE. It's only like two megabits, but it's way more power efficient. I think Humane and Rabbit are both, um, LTE and Wi-Fi, but, um, to get like a wearable you really need Bluetooth.

Host29:34

Just a quick question. Saw the part of the, uh, interface where it had a... Was that a streaming from iPhone? You-

Ethan He29:42

It's a, like a emulated cloud-

Host29:45

Cloud

Ethan He29:45

... project.

Host29:46

With that picture-in-picture, how are you doing the part, portion of the interface where you had the, uh, Google Maps up in the bottom corner?

Ethan He29:51

It's a, it was... It was a... That's on the cloud, so we were streaming that as feedback. I think the ultimate goal is that that disappears entirely, right? That's actually mainly for the demo to like show that it's real and that it's, it's working.

But like I think ideally like it should be totally transparent, that you're just... You have like a personal AI. It can do whatever it needs to do, and it, it will just give you updates on if it, if it, you know, needs more info or, or it's status.

But, um, like it's kind of just the interim solution until we have, I guess... Okay. So sorry, I-

Host30:23

No, no, no, no. It's all good.

Guest 430:27

Yes, it is. Uh, okay, last question, and then we have to move on.

Guest 530:31

Okay. Uh, I, I also, uh, want... I think I've been experimenting with this app, and you have, like, a pockets, uh, here.

Guest 430:38

Yeah, I just put it in my pocket.

Guest 530:39

What, what, what I wanted to say about, about vision is this. Like, I try to experiment with capturing vision and, like, the best solution so far found is, like, like, to buy cheaper Android smartphone and in your-

Guest 430:49

Yeah, yeah, yeah, yeah, yeah

Guest 530:49

... pocket here. Like, under outflows. It's incredible. It has, uh, it has internet connection.

Guest 430:54

Good battery.

Guest 530:55

Good battery. Yeah, it's, it's incredible, and it's extremely cheap, and it's incredible. Like, you have this phone here.

Guest 431:00

Yeah, I, I, I had... I've had, I've had whole, um, whole, uh, whole demos of doing that.

Guest 531:04

Yeah.

Guest 431:04

Because also, you have-

Guest 531:05

Absolutely.

Guest 431:06

Also, it doesn't put people off as much.

Guest 531:08

Yes, yes. And no one, no one even thinks you are recording.

Guest 431:10

Nobody ever says anything.

Guest 531:11

Yes. You just, you just have a...

Guest 431:12

But I, I, I do. I, like... I do think in privacy sense, maybe slightly different than you, I, I do want people to understand. But, like, it, it-

Guest 531:19

Where is the pocket?

Guest 431:19

It was interesting that, uh... Yeah, her style, um, basically.

Guest 531:24

Wow.

Guest 431:24

Um, her style, where... But nobody, nobody, nobody thinks twice about it.

Guest 531:28

It's a different thing. You can just write on it. It's like, "I'm recording it."

Guest 431:30

Yeah, yeah.

Guest 531:31

Like, I'm just saying the convenience of not working the hardware, you can just take the-

Guest 431:35

Sure. Sure, yeah

Guest 531:35

... off-the-shelf hardware.

Guest 431:36

I think... But you, you do need a front, you need a front pocket. So maybe there'll be new AI fashion where it's like, "These are my, my phone pockets." Okay, give it up for him. That was awesome. Uh, you know, you can go up to him.

Um, uh, yeah. Pro... I mean, everyone will, I think, will be sticking around, so you can, uh, obviously go up to him and, uh, get more insight. Awesome.

Guest 632:02

Um, all right. Hey, everyone. Um, one sec. Let me... Oh yeah, it's not connected. Now it is. Nice. So where do I even start? Um, yeah, um, unfortunately, I was kind of not supposed to be here because I'm organizing a brain computer interface hackathon tomorrow, and, uh, I had to, like, somehow get 50 headsets.

Open Source Wearables32:02

Guest 632:25

Um, which is why there will be no presentation, but I will still try to be useful to you as possible. This is the hackathon I was, uh, telling about. Uh, right now, we have, like, lots of people. There will be people from Neuralink.

Um, we'll have, like, 50 different BCI headsets and so on. So if, if you're interested in, like, BCI stuff, et cetera, uh, and if you want to attend the hackathon, scan this QR code and mention that you have been here and...

Oh yeah, my bad. Sorry. And, uh, you might get much more chances to be accepted, uh, because we have, like, 50/50% rate. We try not to accept people who don't have experience. Anyways, um, so yeah, what I will try to help with, I honestly really, really love open source, and I believe, like, all this stuff should be open source.

Which is why now on this short demo, I will just show you all current open source projects, um, and I will try to highlight most important things you need to know about them. And, uh, I'll probably first start with Owl, which, uh, you have just seen, uh, by Ethan.

Uh, so he started that. Um, he... I think he was, like, one of the first people who started, like, open sourcing any kind of wearables. Probably Adam actually was before, but you announced first, so I remember.

Guest 433:36

Yeah. It doesn't matter.

Guest 633:37

But yeah, yeah, yeah. Anyways, so, uh, yeah, this is his repo. Uh, you can, like, check it out. I think I have a QR code here opened as well. If not, just give me a sec. I will just generate it quickly for you.

Uh, all right. Should be... Yeah, just scan this QR code, and you can just access his repository. Uh, yeah. So this is Ethan's. Then there is another one, um, which in my opinion... Well, it's, it's definitely the biggest one.

Uh, it's by the guy who's sitting right there, Adam. I truly believe that Adam is, like, the guy who started all open source hardware movement. So at least I, uh, started doing everything because of Adam, so thank you for that.

Uh, and they have a lot of traction here, 2.6 thousand stars. And if you want to kind of like ramp up your way into open source, um, um, like wearables of any kind, I suggest to start with this repository.

It's probably the biggest one you'll be able to find right now, and the QR code you can scan, I think this one. Yeah, just, like, feel free to scan.

Guest 434:39

I'll send, send one.

Guest 634:40

Yeah, yeah. Cool, cool, cool, cool, cool. Awesome. So they use, uh, I think a Raspberry Pi, also use P32, um, which is kind of technical. You probably don't need this information. But anyways, um, yeah. And now who I am.

Like, a little bit about myself as well, uh, some marketing. So my story starts, uh, very recently, maybe two months ago, after I saw Ethan's and Adam's launch, uh, of, like, their open source hardware stuff. And, uh, I launched my own.

Um, it all started with, um, basically season this is from Humane.

And, uh, we launched, like, just for fun, honestly, Whomane. The idea here was, like, you take a picture and, uh, you scan the person's face, and, uh, we search the entire internet, and we find the person's profile. Um, like, and send it to you via, like, I don't know, like via notification or on the site and so on.

So this is how it looks. Uh, started with that. Had a lot of contributors, was a fun project, but not really... I mean, like, yeah, we don't really want to bring any harm to Humane and so on and so forth.

Just a fun, fun, just fun stuff. So after that, uh, we did... Oh, QR for this, I think we will also send, right?

Guest 435:46

Yeah.

Guest 635:46

Like, uh, cool, cool, cool, cool, cool. Okay. Um, another one I will promote here a little bit is this one. This we launched literally, like, last week. This is pretty much what Adam and, um, uh, and Ethan have, have done, with the only difference that we use right now the lowest power chip.

I think Ethan also uses that. I just try to, like, um, you know, like, uh, as soon as possible to let everyone know that, like, I think it's probably the best opportunity you can currently find on the market.

Uh, it's called Friend. You can check it out. Uh, this is how it looks. Or actually have, have it with me. Um, and, uh, uh, I was supposed to actually show you, like, the live demo as well, but, uh, we just did a very cool update, uh, where we made it work with Android, by the way.

So it's iOS and Android right now. And also- We updated the quality of speech literally three hours ago and made it, like, five times better. So when we launched, I'll be honest, it was, like, completely horrible, and now it's, like, five times better.

It's amazing. Um, and I'm really excited by that. I would want to show you the video, but I know that it's, like, you will not want to have your ears, you know, like-

Harrison36:56

I think I have

Guest 636:56

Oh, yeah? Really? Oh, nice. Okay, let's try. Anyways, yeah, this is the chip that is being used inside of that wearable. Um, and it works pretty much the same as, uh, what Ethan has with.

Harrison37:09

Giving the permissions.

Guest 637:10

Oh, it doesn't, it doesn't work for us.

Harrison37:11

I don't know how to do that. Next. Scan for the device.

Guest 637:16

One sec. Wait, wait, wait. Let me figure out, uh, this... Which one? Christian?

Harrison37:20

Uh, yeah, yeah.

Guest 637:20

Hold on.

Harrison37:21

You can go back a little bit.

Guest 637:22

Okay. Let's try. Oh, nice.

Harrison37:25

And then.

Guest 637:26

Okay, cool. Awesome.

Harrison37:27

So this is started. Giving the permissions here, and hit Next. And scan for the devices here, and then select the device, and that should just connect.

Okay, I think, I think we can just start speaking. Um, yeah, so pretty much testing this for the first time and, uh, let's see if it's able to transcribe what I'm trying to speak. And, uh, yeah, pretty much waiting for the first twenty, thirty seconds to finish, and for it to return us, uh, the output of, uh, the speech from the OpenAI, uh, Whisper endpoint.

And it says, if it's checking... Uh, yeah, there we go. And, uh, yeah, it says-

Guest 638:21

Yeah. So as you have seen, it recorded the speech on the lowest power chip probably ever right now accessible in the market. And, uh, I'm very excited that we actually made it work because, um, it's-- it, it was very hard.

It took like a lot of, a lot of time. Anyways, um, that's pretty much it, I guess. I don't really have anything else for you to show. Uh, this is the final QR code. I know you will have all the links, but this one I really, really advise you to scan because it has basically the collection of like all links in one place.

Um, so yeah, that's pretty much it. Use open source. Uh, I think it's cool, and let's try to build cool stuff together. That's pretty much it. Thank you.

Harrison39:02

Any questions? W- w- did your, did, uh, did your wearable record all that?

Guest 639:08

Uh, no, because-

Harrison39:09

Uh, Harrison, can you come actually and set up?

Guest 639:11

Uh, yeah.

Harrison39:13

You're not, you're not recording right now.

Guest 639:14

Uh, no, no, no. I'm not recording because we launched the update like four hours ago, and I wanted to bring it here. So I brought my thing, and now it's like half broken, unfortunately. But anyways, any questions?

Cool. Yeah, go on.

Harrison39:32

What do you think the biggest next challenge is?

Guest 639:35

Biggest next challenges. Mm. Yeah, I definitely agree with you that the biggest challenge is like editing video and images and so on. It's like very hard as hell. And I think to make software useful as well is very, very hard.

Um, like, yeah, we can all do maybe like recording from the device and like attach maybe some battery and so on. But like how to make it actually like sexy, that's, that's hard. Like how to make it, you know, do actions, how to make it remember everything and so on, so forth.

That's the biggest challenge. So yeah. But I agree with everything you said.

Harrison40:06

Lemme...

Guest 640:06

Huh?

Harrison40:07

I, I would add that I think the biggest challenge is making people want to wear it.

Guest 640:11

Oh, yeah.

Harrison40:11

And that are not the te- the tech bros or, or girls of Silicon Valley.

Guest 640:16

Yeah. Adam, uh, the guy who created the biggest open source, uh, thing, he said that, uh, the biggest challenge is to make people want it, basically, right? Am I saying... Yes. Um, so that's, that's, that's one of his suggestions.

Um, yeah, go on.

Harrison40:30

Uh, what's been the challenge in reducing latency?

Guest 640:33

What's been the challenge of reducing latency? Um, honestly, it was just like software issue, um, because like this chip is like not that widely used. There is not so much documentation, not so many projects and so on. So just like a matter of like trying a lot of things.

And also it doesn't have huge onboard memory. So like how do you store very quality memory on a very small, uh, very quality audio on a small memory chip, and then send it to phones extremely basically non-stop. That was pretty hard, but we solved it like five hours ago, and, uh, that's pretty much it.

So... Yeah, go on.

Harrison41:09

How did you improve the quality of voice by that much?

Guest 641:12

Um, we just made it work. Like we used, um... We had four thousand, uh, pretty technical, I think. I don't know if you want this, but anyways. We used four thousand hertz, uh, like quality. It was like super bad because the memory was too small, and now we just found a way to like compress it and improve it to like sixteen thousand, which is like pretty great, which you have heard on the, um, on the video.

So it can like recognize pretty much anything, even like multiple people speaking, even if you will be there and the device will be here. So yeah. Anything else?

Harrison41:46

I think we can leave the-

Guest 641:47

The last one. Go on.

Harrison41:49

Okay. Pretty like how far-

Guest 641:53

How far it can detect-- how far the device can detect audio. Correct? Cool. So good one will be like, you know, two feet from, from person to person, like good. Uh, if you have maybe like four feet from each other, the quality will be like fifty percent accuracy.

So yeah. Cool. Thank you all. Do open source.

Memory & Personalization42:22

Harrison42:22

Cool. Cool. So what I wanna talk about is a lot less cool than all this hardware stuff, so I feel a little bit out of place. Um, but my name is Harrison, uh, uh, a co-founder of LangChain. Uh, we build kind of like developer tools to make it easy to build LLM applications.

One of the big parts of LLM applications that we're really excited about- ... is the concept of kind of like memory, um, and personalization. And I think it's really important for personal AI because, you know, hopefully these, uh, assistants that we're building remember things about us, and remember what we're doing, and things like that.

Um, we do absolutely nothing with hardware. So, uh, when we're exploring this, we are not building kind of like hardware devices. So we took a more kind of like software approach. Um, I think one of the use cases where this type of like personalization and memory is really important is in kind of like a journaling app.

Um, I think, uh, for obvious reasons, when you journal, you expect it to be kind of like a personal experience where you're kind of like sharing your goals and learnings. And I think in a scenario like that, it's really, really important for, uh, if there is an interactive experience for the LLM that you're interacting with to really remember a lot of, of what you're talking about.

So, um, this is something we launched, I think, last week. Um, and it's basically a journaling app with some interactive experience, and we're using it as a way to kind of like test out some of the memory functionality that we're working on.

So I, I, I wanna give a quick walkthrough of this, and then maybe just share some high-level thoughts on, on memory, um, for LLM systems in general. Um, so the UX for this that we decided on was you would open up kind of like a new journal, um, and then you'd basically write a journal entry.

Um, and I think, uh, th- this is kind of like a little cheat mode as well, 'cause I think, um, this will encourage people to, to say more interesting things. So I think if you're just taking like a regular chatbot, there's a lot of like, "Hey.

Hi. What's up?" Things like that, and I don't think that's actually interesting to try to like remember things about. Um, I think it's more interesting if you talk about, uh, personal things. And so let me, let me try this out.

Um, I'm giving the talk right now. And then I can submit this, and then the UX that we have is that a little chat with a companion will open up. Um, okay, so yeah. So right before this, I told it that I was about to give a cha- about a, a journaling app.

And so it, uh, kind of like remembered that I was going to, um, uh, do all that. Um, is there a particular part I'm most excited to share? The memory bit. I don't know. So this is, this is, this actually worked on the first try, so I was a bit surprised by that.

So um, that's good. Um, and, oh, okay, so how do you plan to tie in your love for sports with the theme of memory during your talk? So, uh, before, when I was talking to it, I had mentioned that one of the things that I wanted to talk about, uh, was, you know, how a journal app should remember that I like sports.

So I guess it remembered that fact as well. Um, amazing. So I can end the session. So the, so the basic idea there... And again, this is, you know, we're, we're not building this as a real application. We would love to enable other people who are building applications like this.

Um, I think the thing that we're interested in is really, like what is, what, what does memory look like for applications like this? Um, and I think you can see a little bit of that if you click on this memory tab here.

Um, we, we have like a user profile, um, where we basically kind of like show what we learned about, uh, a person over time. And then we also have a more like semantic thing as well. So I actually could search in things like Europe, um, and I'm going to Europe, uh, kind of like after our wedding.

I love Italy. Um, and, and so basically there's a few different forms of memory. Um, and, uh, if, if you'll allow me two minutes of kind of just theorizing about memory, um, we're doing a hackathon tomorrow, and, uh, maybe some of you are going to that.

Um, Swix signed up. I don't know if he's actually- Yeah ... gonna show up. In the hack, in the hack. Um, so, uh, very quickly, like how, how... I think memory is really, really interesting. Um, it's also kind of like really vague.

At a high level, I think like there's some state that you're tracking and then... But how do you update that state, and how do you use that state? These are like really kind of like, uh, vague things, and there's a bunch of different forms that it could take.

Um, some examples of, uh, kind of like... Yeah. Thanks. Uh, some examples of, of memory like that, that a bunch of, uh, real apps do right now, like conversational memory is a very simple but obvious form of memory.

Like if you're... It remembers the previous message that you sent, like that is like incredibly basic. But I would argue that it can, can fall into this idea of like how is it updated? What's the state? How is it combined?

Semantic memory is another kind of similar one, um, where it's, it's a pretty simple idea. You take all the memory bits, you throw them into a vector store, and then you fetch the most relevant ones. And then, uh, I think s- uh, one of the early types of memory that we had in LangChain was this knowledge graph memory, where you kind of like construct a knowledge graph over time, which is maybe like overly complex for some kind of like use cases, but really interesting to think about.

Um, so LangMem, like name TBD- ... um, is, uh, some of the memory things that we're working on. And we kind of wanted to constrain how we're thinking about memory, uh, to make it more tractable. Um, so we're focusing on like chat experiences, um, and chat data.

Um, we're primarily focused on like one human to one AI conversations. Um, and, and we thought that flexibility in defining like memory schemas and instructions was really important. Like one of the things we noticed when talking to a bunch of people was, like the exact memory that their bot cared about was different based on their application.

If they were building like a SQL bot, that type of memory would be very different from the journaling app, for example. Um, so there's a few different like memory types that we're thinking about. All of these are very like early on.

Um, uh, I think one interesting one is like a thread-level memory. An example of this would just be like a summary of a conversation. You could then use this, uh... You could extract like follow-up items. And then in the journaling app, you could kind of like follow up with that in the next conversation.

We actually might have added that. That might be why it's so good at remembering what I talked about in the f- previous talk. I forget. Um, uh, another one is this concept of a user profile. It's basically some JSON schema that you can, uh, kind of like, uh, update over time.

Um, th-this is, uh, this is one of the newer ones we've added, which is basically, like you might want to extract like a, a list of things. Um, similarly, like define a schema, extract it, but it's, it's kinda like append-only.

Um, so an example could be like restaurants that I've mentioned. Like maybe you wanna extract the name of the restaurant and what city it's in. Um, if you're kind of like overwriting-- if you put that as part of the user profile and you overwrite that every time, that's a bit tedious, so this is like append-only.

Um, and then we do some stuff with like knowledge triplets as well, and that's kind of like the semantic bit. Um, I think, I think probably the most interesting thing for both of these is maybe like how, how it's like fetched.

Um, so I don't know if people are familiar with the Generative Agents paper, um, that came out of Stanford last summer-ish, but I, I think one of the interesting things they had was this idea of fetching memories, not only based on semantic-ness, but also based on recency and also based on like importance.

And they'd use an LLM to like assign an importance score. Um, and I, I thought that was really, really novel and, and enjoyed that a lot. Um, and yeah, that's basically it. So yeah.

Ethan He49:34

Tons of questions.

Harrison49:35

Questions.

Ethan He49:35

Well, yeah. Yeah. I think they all about all the same things, and, uh, I think a lot of your approach makes a lot of sense. Same, same kind of compromises you have to make for simplicity. But like to give it true knowledge, and you talk about like the triplets and like, um, like how do you think that we can get to the point where we can have more of a dense graph rather than just simple propositions about you?

Because like our memory works and it's all relational to like, you know, to people, to places, and that, that's actually important information. And so like, do you think we'll be able to figure out a simple way to do that, or is it just gonna be too hard?

Harrison50:10

Yeah. I-- the honest answer is I don't know. Um, I think even today, like if you had that, and I think like-- well, there's, there's two things. Like one's like constructing that graph.

Ethan He50:19

Yeah.

Harrison50:19

But then the other one's like using that in, in generation. And like even today, like for most like RAG situations, combining-- if you combine a knowledge graph, that's often not taking advantage of like bot. It, it's, it's really like, it, it's, it's...

There's different ways to put a knowledge graph into RAG and it's not, uh... Yeah. It, it's very exploratory there, I'd say. So I'd say like one issue is just even creating that, and then the other one is like, um, using that in RAG.

So I think that's like a huge... Yeah. I don't, I don't know.

Ethan He50:49

Yeah. We'll see. It's interesting.

Guest 750:56

Um, yeah. Uh,

Ethan He50:57

Questions?

Guest 754:25

Yeah. So when I see these systems, and I guess that's a question pretty much for everyone who presented, right? Like when I think about these systems, I always think, what does ten years of memory look like? And- And a lot of the facts that we remember are either not relevant anymore or probably false.

So how do you think about, like, memory decay?

Harrison54:48

Yeah. I, I think there absolutely needs to be some sort of memory decay or some sort of, like, invalidating previous memories. I think it can come in a few forms. So like with the generative agents kind of like, uh, paper, um, I think they tackled this by having some sort of recency, uh, weighting and then also some sort of importance weighting.

So like, it doesn't matter, like, you know, how long ago it was, there's some memories that I should always remember, right? Um, but then otherwise, like, you know, I, I, I remember what I had for breakfast this morning.

I don't remember what I had for breakfast like ten, like... And, and maybe I should, but I don't think that's, like, important. Um, so yeah, I think like recency weighting and importance weighting are two really interesting things in the generative AI or the generative agents paper.

Um, another really interesting paper with a very different approach is MemGPT. Um, so MemGPT uses a language model to, like, actively kind of like construct memory. So like in the flow of a conversation, like the, the agent basically will decide whether it should write to like short-term memory or long-term memory.

And, um, I, I think that's a... Yeah, that's a... It, it... I think it's actually quite a different approach, because I think in one you're having the application, like, actively write to and read from memory. And then in the other one, the one that we're building, it's more in the background.

And I think there's pros and cons to both. Um, but I think with that approach you could potentially have some, like, overwriting of memory or, or, yeah.

Host56:10

Cool. One last question.

Guest 856:11

Last question. Also, that generative agent small build paper, amazing.

Harrison56:15

Yeah.

Guest 856:15

Like one page, they have like an exponential time decay. Every day stuff is less relevant. Definitely recommend.

Guest 356:23

Hey, Harrison. Thank you for presentation. I just want to ask about, there was a new paper called Rapture, I think, and it feels like it's a really cool approach to memory because sometimes when you want to say something like, "Who am I?"

Or, "Roast me," it's really hard to do with RAG and these type of approaches, but the Rapture could be a nice way to tackle that. What's your opinion on that?

Host56:47

Can you summarize the paper?

Guest 356:49

Uh, I think it's about like, uh, doing like partial summarization and in like a tree form, and we are trying to experiment with that, and it seems like... But I think you know more about it. I just like found out about a couple days ago, so.

Harrison57:03

Yeah. No, so I think the idea is basically for RAG. You, you chunk everything into really small chunks, um, and then cluster them, and then basically hierarchically summarize them. And then, uh, when navigating it, you can go down the different nodes.

I hadn't actually thought of applying it to memory, but that actually makes a ton of sense. So one of the issues that we haven't really tackled is in this journaling app, if you notice in here, there's, um, a bunch of ones that are really similar, right?

Um, and so like there's a clear kind of like consolidation or update or something procedure that, that kind of needs to happen that we haven't built in yet. And so I actually love the idea of doing this kind of like hierarchical co-- summarization and, and, um, maybe like collapsing a bunch of these into one, and maybe that runs as some sort of background process that you run every, um, I don't know, you, yeah, you run every day, week, whatever.

It collapses them. It accounts for recency to, to account for the issue that was brought up earlier around wanting to like maybe overwrite things. Yeah, I think that's a re-- I had not thought of it at all, but I think that's really, really interesting.

Guest 858:03

That's kind of like sleeping.

Harrison58:05

Sorry?

Guest 858:05

Like when the brain consolidates-

Harrison58:07

Yeah.

Guest 858:07

-memory during sleep.

Harrison58:08

Yeah, yeah, yeah. Yeah.

Wrap-Up58:10

Host58:10

Uh, cool. I think I'm, I gotta wrap it here. Thank you so much, Harrison. A round of applause for him.

Um, so I think there's a real reason why this is not just a hardware-only meetup. It started as a wearables meetup, but then we added full-time voice, and then we added memory, um, because it's all components of building, uh, your personal AI.

Uh, so we have fifteen minutes until we have to clear out of here. Uh, everyone is-- looks like you're sticking around to just chat. If you want to just see the devices and talk to Harrison, go ahead. Be my guest.

Um, there are like three hackathons spawning from this thing, so- Uh, I'll send all the links on Puma. Thank you. Thank you very much.

Harrison58:47

Thank you.

Host58:53

Thank you.