LALatent SpaceMar 27, 2025· 48:59

Building Manus AI (first ever Manus Meetup)

Tao Zhang, co-founder of Manus, explains how Manus gives LLMs 'hands' to take real-world actions, inspired by MIT's 'mens et manus' (mind and hand). Manus assigns each task its own cloud virtual machine via E2B, provides pre-paid data APIs (stock, social media), and a knowledge system to remember user preferences. Their previous product, Monica.im, grew to 20M monthly active users and $50M ARR before they pivoted from a failed AI browser project to Manus. On the GAIA benchmark, Manus achieves a per-task cost of ~$2, far cheaper than OpenAI's Deep Research (~$20) and prior SOTA. Zhang emphasizes a 'less structure, more intelligence' philosophy, opting not to predefine workflows but to give the model rich context and tools. He also discusses plans to partner with Cloudflare, pay paywalls on behalf of users, and integrate more APIs based on usage patterns, while ruling out building their own foundation model.

  1. 0:00Mens et Manus
  2. 3:27Gaia Benchmark
  3. 13:01Why Manus?
  4. 20:01AI Browser
  5. 28:00Cursor Inspiration
  6. 31:23Building Manus
  7. 37:46Less Structure
  8. 40:17Quality & Context
  9. 44:27No Own Model
  10. 45:55Data Access

Powered by PodHood

Transcript

Mens et Manus0:00

Tao Zhang0:00

The first slide is, uh, we're gonna answer this question. It's like, uh, uh, what is Manus? Uh, this name, Manus. Yeah. Actually, this name, this word comes from an old Latin word, yeah, uh, which is also the MIT's motto.

It's like, uh, mens et manus, which is men and hand. Yeah. It's like people, yeah, just, kind of, only think. You have to take actions. So from our perspective, we think for the past two years, the LLM is already taking over the whole world.

And, um, that's not, not with the latest models. Maybe even like one year ago, we think the LLM models are already powerful enough and smarter than most of humans, uh, in every vertical. But we think LLMs is only like the men.

It can only think, only reasoning things as ha- as that. But if you want to take real impact into the physical world, into real life, you have to build hands for the AI system, not only men. Yeah. If you only think, think, think, you can't do anything.

Like for me, actually, right now I'm like, uh, thirty-eight this year, and I started coding from like, uh, eight years old. Yeah. That's thirty years ago. So my first program-programming language is, uh, the Logo programming. I actually I don't know if, if anyone here knows that.

The Logo is a script, uh, langwa-- programming language for drawing, for, for drawing things. Yeah. So at that time, yeah, thirty years ago, back in China, uh, not everyone has, has his own computer. So I only have like two times a week, I get the chance to go to the computer room in my primary school.

So besides that, besides that two, two, two, two chances, I only have to write all the scripts, all the code on my notebook. Yeah. That is not this notebook, you know. It's a, it's a physical notebook. So even I, I'm already like maybe...

I, I, I think I, I might be the first or second here in our coding lessons thirty years ago. But even, even that, I, I, I just can't put the codes right without the computer. You know, if you only think the code in your, in your mind, in your brain, without real debugging and testing on real computer, you can't write the code just in one, uh, write the code right just in once.

So that's also happened in the LLM world. We think for the past two years, the problem is that we've already have, have a very powerful LLM models, but we just lock it in a black room or gi-give it only the pen and the, and the, and a, a very little notebook.

But we don't give it computer. We don't give it outside world access. And we ask them to do very hard work, but they can't do that. So Manus is kind of like a... We are-- We want to make hands for LLMs to help the AI to make real impact to the world.

That's how our name Manus comes from. And also I'm going to Boston next week to visit MIT. That's, that's my this trip's purpose. Yeah. We must bring the name to MIT campus. Yeah. So, uh, during building Manus, we are also doing some benchmark.

Gaia Benchmark3:27

Tao Zhang3:27

Yeah. We're also doing some benchmark. That's what you see in our video on, on our website. And it's just an early checkpoint at the end of January. So actually right now we already have a new, we have a new result for Ben-- uh, for, for the Gaia benchmark.

And, uh, we got a lo-- we got a, a, a big improvement, uh, uh, which I just saw yesterday from my colleagues. Uh, they've already, uh, run the benchmark again, and right now we are far better, uh, than this one.

This one is at the end of January. But you can say that, uh, during OpenAI release their Deep Research at, uh, at the first of, uh, like, uh, like February, yeah, they, they choose Gaia benchmark. But I don't know if anyone tonight, uh, familiar with the, uh, Gaia benchmark.

Yeah. But, uh, later I will show some cases for you to, to understand why OpenAI research and also Manus we choose this benchmark, uh, to evaluate our system. Yeah. But that's what we achieve here. Yeah. You know, as a company who, who, who, who's not trained models and only use APIs, uh, we, we, we think that is already a good enough job for us.

Yeah. And we're not just have some, like, best performance. We are also like the most cost efficiency. Yeah. We found some articles which discuss about the previous SOTA and OpenAI Deep Research about when they run Gaia benchmark, uh, uh, what's the price for each tasks.

It seems like in that vertical, uh, it says like maybe twenty dollar per task for the previous SOTA and, uh, even higher for, uh, OpenAI's Deep Research. So during our test, uh, we just solved the Gaia task with like, uh, it has like one hundred more tasks.

We solved it at a average cost at two dollars. So it's like, uh, two times cheaper than OpenAI and five times cheaper than previous SOTA. Yeah. So let's just, uh, give, give you guys some quick understanding of what is Gaia's benchmark.

Yeah. I, I, I can... Yeah, I continue. Yeah. So like th-this is the, uh, this is the, the, the, the, the questions from the Gaia benchmarks task. Yeah. It's like, uh, uh, you know, asking you, uh, there is a, a picture on NASA's website on like, uh, two thousand six and January twenty one.

So the agent have to go to the NASA's website to find the exact image and the two... There are two astronauts, uh, the one with one appearing much smaller than the other one. So the agent have to go to the website, find the image and identify the smaller one.

And then they said, uh, "What's the name for the astronauts?" And asking, and the, the, the final question is, "How many minutes did he spend in space?" Actually, on the whole internet, there is no exact answer for this question because this astronaut go to space many times.

Yeah. So Manus has to like go over the-Internet to search all the, uh, space mission for this astronaut and add, add all the space time together and to give the final answer. Yeah. So Gaia, Gaia Benchmark task is more like this.

It's kind of simple if we are real human. Yeah, we just can... We, we gotta spend some time to, to know how the-- how, what, what is the task and use Google, use everything when, and, uh, with much of time.

And eventually, uh, for, for, for an average human, we can solve these tasks. But you know, like two years ago, maybe one year ago, it's really hard for a normal AI, like ChatGPT, chatbot things, to solve these tasks.

Yeah, but you know, right now, uh, with Manus, it's very easy. Yeah. You can see Manus just, uh, uh, first has, has its planned and decided to do some search and find the, uh, NASA's, NASA's, uh, NASA's image and the name of the astronaut and find all the space missions, uh, he did and, uh, combine all the space time together, then deliver the final, final result.

Yeah. That is one. Yeah. And another one, and it's kind of weird, you know, yeah, because I am the person, uh, I'm co-founder of Manus, but I'm the product, I'm the product guy. Yeah. I'm kind of product. Yeah.

But you know, since we're a very high startup, uh, we are very-- our job are very s-flexible. So the Gaia Benchmarks, all these tasks is actually run by me. Yeah. So when the first time I saw this task in Gaia Benchmark, I was kind of, "What?"

Yeah. It is, it, it, it's like, it's like a task from, uh, World of Warcraft. Uh, so it, it, it takes me back to twenty years ago. Yeah. Just, uh, what's the, what's the best, uh, uh, group when you want to, when you want to take, take some whole boss down.

Yeah. So Gaia Benchmark is very interesting. Yeah. It's full of task, uh, like this. And also, also for, also for this one. Yeah. Also for this one. Yeah. It gives a picture like this and, uh, ask, "What is the brand of this?"

You know, Gaia Benchmark al-also teach me something because before this task actually, I don't know what is... What it is in English. Yeah. But, but, but, but the Gaia Benchmark just, just, uh, teach me, uh, this word, th-this word.

You have to find the brand the dogs are, are wearing and, uh, uh, in fact, and, uh, what meat is mentioned in the story added, uh, on the, um, on, um, the ambassador's blog on this exact date. So, you know, this, th-this, this task, uh, is kind of challenging because you have to, uh, really understand the...

What, what's in this image and then search on internet to determine the brand of, of, of this one. And go to its official website and find the exact date. And the, uh, that, that actually is really long. So Manus has to, uh, scroll, scroll a lot.

Yeah. We can, we, we can check it here. Yeah. Yeah. You, you... Manus is, uh, is scrolling down, scrolling down, and find that article in that exact date and decide to open that, that article. And, uh, scrolling down, scrolling down, scrolling down.

Uh, later. Yeah.

Oh, yeah. It should be, it should be this article. Uh, but this article, if you, if you open it, uh, it's, it's very long article. So Manus has to scroll it. I, I remember the answer is at the, like, uh, like maybe two, uh, uh, like two third, uh, at the, at the position of that, of that article.

Yeah. So, you know, Gaia Benchmark is, is full of tasks like, like this. And it's really... And maybe I can say it's very easy. Yeah. But even as a human, you have to do a lot of work. So before, before this, many agents, uh, can't, uh, deal with the tasks.

But right now with Manus, actually, we've already hit maybe, maybe after six months, we're gonna hit the limit of Gaia Benchmarks. So we are also finding some very interesting internal evaluations of our systems right now. E-- For example, uh, during building Manus, we were already, uh, using some task from Fiverr and Upwork.

Yeah. Which, which people will, will, will pay. So we have, uh, some like internal evaluations, uh, from Fiverr and, uh, uh, and, uh, yeah. And, uh, another thing I want to, I want to say here is like, uh, because in our video we are targeting Manus as like the first general AI agent.

Yeah. And, uh, why we feel so confident that we can say we're general? You know, because, you know, during building Manus, uh, one of my intern, uh, he goes to the YC's website and found that during YC's W twenty-five batch, there are like, uh, twenty-one, uh, projects is like agent-oriented.

Yeah. So we look over these twenty-one, uh, projects. And since... I, I don't know if all you guys know, uh, some projects from YC, uh, they don't even have a real demo. They only have a video on their website.

So if they have demo, we will, we will run the demo and compare it to the results of Manus. If they don't have demo, we can only just compare the use cases they propose on in their video. And finally, uh, we will- We're very lucky to find that actually Manus covers like, uh, seventy-six percent of that batch's agents.

Uh, they are all different kinds of ba-- of, of, of agents. Some for medical, some for legal, and some for, uh, marketing, whatever, uh, different kinds of... And in most of scenarios, actually, uh, but, you know, it, it, it's very like, like, like biased because I, I built Manus.

We think in, in most scenarios, Manus' output is actually better than the virtual agents. So that's why we target it as the first general agent. By covers you mean implement? Oh, yeah, yeah. It's like we can perform their- You mean you can implement.

Yeah, yeah, yeah, yeah. Okay. Yeah, so that's one is Manus. And the second, uh, part I want to cover is, uh, why? Yeah, you know, yeah, why we build, why we build Manus? Yeah. So, uh, if we talk about Manus, actually, you know, this company starts like two years ago, uh, two years and four months ago.

Why Manus?13:01

Tao Zhang13:21

So Manus is actually our second product. We have our first product. Yeah, our first product, uh, we called it Monica.im. Yeah, so many of you may be curious about why do we choose this domain name, Manus.im. It's because our, our first product we used the im domain name.

It's like we... First product is Monica.im. And, uh, I know, yeah, not, not, not all of you know what Mon-Monica.im is. Uh, actually, Monica.im is the, uh, a Chrome browser extension, and this project has started one week before ChatGPT.

And why we want to build Monica is because, you know, when people are using ChatGPT and Azure Open AI service cloud, there are products that you have to switch between apps, copy-paste text, copy-paste the images, right? It's very frustrating.

So we think we, we, we should let users to like use LLM just in the context without switching. So to help you better understand what we built before, yeah, I can just show you some, some features of Ma-- of, of Monica to you.

Yeah. Uh, like the transformer circuits. This is the best blog in AI industry, I think. Yeah. I, I read every piece of, of, of them from Andrew Ng. But sometimes, you know, as I'm, I'm not a native speaker for English.

Yeah, so sometimes when I read a very long research article, I just get kind of tired. So right here, yeah, uh, this sidebar we-- is, is what the product we built. It's the, uh, Monica. Uh, and we have a very little feature.

We call it simplify article. Yeah, so it's like when you read a long article, you can just click the simplify article, and the whole article will like, you know, the, the, the long section will, will, will be, will be, will be simplified.

And why we want to build this feature is because, you know, uh, there are a lot of AI, AI apps or they have a function like a page summary or summary article. But actually for me, I think page summary doesn't work because, you know, uh, it...

Because all these page summary functions, they don't know your attention. Yeah. When it do the summary thing, it will, it will lose focus, and it, it doesn't understand your attention. So you, you, you will be kind of like, like cautious about, okay, it, it, it, it's short, but did it lose any information?

So we decide to build this feature, simplify article. Yeah, you can just, uh, keep scrolling this one, but without losing all the like tables, images inside it. And if some paragraphs you want to digger, uh, dig around, you can just move your mouse here, and you can see the original one.

So that's, that's a very one little feature. And we also have other features like in, in YouTube. You can, uh, right here if you use the Monica, uh, there, there will be a little bracket here. You can generate a summary, and even you can generate a podcast based on this YouTube video and listen on your phone.

Yeah. So that's in YouTube. And also, uh, because, uh, we are in AI industry, we read papers every day on Arxiv. So you can gen... On Arxiv, you can just summarize the, uh, uh, or expand this paper just inside the Arxiv.

So that's what Monica do. Actually, we, we... I just introduced three features to you. Actually, we have, we have hundreds of them. Uh, like we have one in the Gmail. You can just, uh, uh, use ChatGPT, have draft, draft a, a better email inside the Gmail.

One year before Google implemented this feature themselves. Yeah. So it's like when you are using Monica, uh, you can just, uh, uh, leverage the power of LLMs just inside your browser. Yeah, that's what we built, uh, like two years ago.

And, uh, actually, uh, the Monica.im just went very well, and now it's the best in its category. Yeah. We have like, uh, twenty million monthly active user, and also generating like fifty million ARR right now, and it's still growing.

Uh, we are targeting at maybe thirty million ARR this year. Yeah. So that's the first thing we built. Yeah. Monica.im. Yeah. We are not a, we are not a new standard. We are in this industry one week before ChatGPT.

Yeah. But you know, uh, when I say Monica.im is a Chrome browser extension, it creates... It creates three problems. And the limit, the ceiling for us is because not everyone is using Chrome, and not everyone is using browser, and not...

And actually, you know, when we talk about extension, many people didn't know what extension is. So after building Monica, we are always come... We are always thinking about- Next move. Uh, what is the next ChatGPT moment? Because all you guys know, uh, GPT-3 is far, uh, uh, is, is far earlier than ChatGPT, right?

Yeah, actually, GPT-3 is actually very good. But before ChatGPT, uh, only a few startups found the power, the potential of GPT-3. Uh, as I can remember, it's like Jasper, uh, Copilot AI. Yeah, the on-only a few startups found the, the power.

It's because ChatGPT just, uh, uh, introduced the, uh, a, a paradigm of how to interact with LLMs. Yeah. Yeah, ChatGPT defined how humans will interact with LLMs. But we, we think chatbot is not the final form. Yeah, because at the very, at the ve-very bottom of my heart, actually, I think the ability to ask, yeah, especially to, to ask the right question, is very hard for, for most human.

It, it is very hard. Yeah. And, uh, with, with, with, with this, with this assumption, it's like it just limits the potential of LLMs. Because, you know, uh, it, it-- LLM is very smart, but if your product, your chatbot, its product value is limited by your users.

Yeah, they a- they ask the right question, and, uh, they will feel the value. If they can't ask the question or they can't ask the right question, they can't feel the value, then it's very hard for you to scale your product.

Yeah. So we think that, that's, that's a problem. Yeah. And as we are already very successful in browser extension, so it's very easy for us, uh, to think, yeah, maybe our next product should be an AI browser. Yeah.

AI Browser20:01

Tao Zhang20:01

You know, yeah. We think, yeah, now everyone's using browser extension, so maybe we should build the browser ourselves. Yeah. So that's our original thinking, like one year ago, uh, last March. Yeah, just one year ago, last March. And, uh, actually, we are taking very serious steps in the AI browser.

Yeah. I, I, I think nowadays you heard a lot of concepts about AI browser, but we started to build AI browser one year ago, and with very serious, uh, efforts. At that time, uh, we only had like forty people, and we put twenty of them into this project.

Uh, as a startup, that, that's a lot, uh, investments. Uh, and we build AI everywhere, uh, inside the browser, and we also post-trained a small side model, yeah, especially for browser. And, uh, we build, uh, right now you, you've already know the project browser use.

Actually, one year ago, we, we build our own browser use even before browser use. And our browser can interact with other apps. Like, you can just export a page into Excel. Yeah, things like that. We've already built that, and it's working.

It's working. Not, not a working prototype. Actually, it's, uh, it, it's a real browser. Uh, we, we built it. Uh, we have this dashboard which can, like, uh, uh, generate these cards, uh, based on the topics of all your tabs, so you don't have to organize them all by your- by yourself.

And you can just, uh, uh, uh, up-scale any images on your webpage, videos. Yeah, and you can just preview summary, any links on webpage. Yeah, we, we built a lot. And also it can automate, yeah, things like this.

Uh, on LinkedIn, you want to search some, like, iOS design experience and, uh, uh, the, the, the browser will, like, go over all the candidates and to see if they have this experience, uh, on, in this page and even in the next page.

Yeah. So, you know, we, we, we, we, we take, we take a lot of resources on that project. And, uh, from March to September, we-- our team spent six months in the, in the AI browser project. Yeah. But finally, uh, last, I think last September, we, we found there are some problems.

Yeah. The first problem at least in here is that actually browser is for single user usage. Yeah. Hey, Parker, uh, can you give me some water? Thank you. Water. Yeah. We think browser is for single user. Yeah. So after the AI control your browser, we find this very frustrating because, you know, if you are only, uh, looking at some videos on YouTube that AI controlling browser, you will think, "Oh, wow, Sam, it's cool."

But if it's you uh, who is sitting, uh, uh, sitting in front of your, your computer, you will find it very frustrating because once the AI started controlling your browser, you have to take your hands off the keyboard and the mouse.

Even one very small move will just, uh, break the whole process. And another, another problem is, uh, because, uh, the-- you, you, you don't-- you have no idea when AI will finish its job. So when the AI starts its work, you have to keep your attention to it.

And, and imagine that. You can't touch it, but you can't, you can't, you can't go away from it. You have to look at, but you can't touch it, and you can't switch to other apps to continue your work.

And so after six, after six months, we, we, we built that, we decided to sunset the project. Uh, just two weeks before the original date we, we want to release it. Yeah. We, we never re-re-release the browser, yeah, after six months of work.

And you know, that is a very dramatic moment for us. It's like, uh, uh, the day we decided to sunset the, our AI browser project, and after we, we make the decision, I take the flight back to Beijing.

Yeah. And after the flight landed, I, I, I opened my, my phone and, uh, to, to scroll the Twitter. And , uh, the, the, the first thing I saw is Josh Miller's video. I don't know if anyone of you know Josh Miller, the browser company who made the Arc.

Yeah, Arc Browser, you know. Because, uh, when we start, when we start, start the browser project, we got a lot of inspirations from Arc. So actually, we, we, we take real respect for the team and even for, and for, and for Josh Miller.

Yeah, so the first thing I see Josh Miller's video, and he-- In that video, Josh Miller announces that they will discontinue Arc and move to the next project. I was like, "Oh, okay, okay. Yeah. Now we're on the, now we're on the same page."

Yeah. So, uh, th-that's our, like, maybe not very successful attempt to our second product. Yeah. But, you know, even we have some problems, but we also have some learnings of, uh, during building the AI browser. The learning is that actually we think AI should use browser because, you know, during, uh, we building the AI browser, we found out that AI is actually very good at controlling browser because AI knows a lot of techniques humans doesn't know, or we only know one trick, but AI knows it all.

Like, uh, there, there, there is one test, uh, during the, uh, Manus' building is like, uh, one task ask the Manus to show some video on YouTube and, uh, yeah, yeah, that, that's just the task. So some video on YouTube.

And when Manus go to YouTube, it press three. When, when, when... Because I, I'm the one running that task. Actually I'm very like, like, "What? What? Why? Why? Why do you press three?" Yeah. I don't know if anyone here tonight knows on YouTube if you press three, what happen.

Three? Yeah. You know, tonight we have like twenty or thirty people here. Nobody know this trick. Yeah. On YouTube, there's a shortcut key, uh, from zero to nine. When you press it, it will jump- Oh, yeah ... the percentage.

You know, if you press three, you will jump to the thirty percent of the video. That-that's just one example, you know. But, but AI knows so many tricks, not just for YouTube. There are so many websites on the, on the internet.

So it's like we think, "Okay, AI should use browser." It knows a lot. But the purpose is will not use your browser. AI should use its own browser so you can continue your own work. So that's the first learning.

And the second learning is that the AI should be in the cloud. Yeah, the browser should be in the cloud. Then you don't have to waste your attention to it. Yeah, AI just continue to work, uh, in the cloud, and once the job is done, it notifies you.

And that's okay. Yeah. And the third learning is hard. It's really hard. It take us six months to get this learning cycle. Instead of convincing people switching browsers, because browser is really a, like, very huge project, and Chrome, Google, they invest a lot of resource on that for, for years.

And if you want to convince people like, "Okay, we have a cool new AI browser. You should try this." But, you know, before the user can feel the AI power, he will ask so many questions. "Why, why don't you have this feature from Chrome?"

Yeah. "I, I use it every day. I want, I want that." Yeah. People will, you know, people are so familiar with, with, with browser. So if you are asking them to, to switch browser, and they will ask you to build a, a very powerful browser before AI features, and that makes a lot of problem as a startup, you know?

Yeah. So that's three learnings we learned. Yeah. And while we were working on browser project, actually the world is changing. As I just mentioned, we work on AI browser from March to September last year, and I think all of you, all of you can remember last July, Cursor, right?

Cursor Inspiration28:00

Tao Zhang28:21

Cursor just got so popular around the world. Yeah, that's a thing. Yeah, that's a thing. And actually, at the very beginning of Manus project, Cursor just inspired us a lot. Yeah, a lot. Really a lot. Why is that?

Because, you know, um, Cursor's interface is some kind of like this. Uh, I don't know if you use Cursor. This is just so famous, and this is the, uh, Bay Area, San Francisco. I assume everyone would know what Cursor is.

Yeah. So do we use Cursor in our team? Uh, me as the product guy and our chief scientist, Pete, who is in the video, and our CEO, Red, we three all know, we all know coding. Yeah, we, we, we, we, we, we, we, we, we start coding like twenty or thirty years ago.

And, uh, we are, we are familiar with coding since. But you know what's the interesting part is that when we look at our colleagues in our company who doesn't know coding, who doesn't know coding, they are also u-use Cursor too.

And when you watch these people use Cursor, that's the most interesting part, is that when they are using Cursor, they don't actually care about the left side. It's like whatever code, uh, Cursor writes, they just accept, accept, accept because they don't know coding.

So, so they can't judge that, that Cursor gives the right code. They only care about the right side. They just keeps click, uh, accept, accept, accept, accept, accept. And the tasks our colleagues are using with Cursor is also amazing because when we, like us coders, we are using Cursor, we want to use Cursor to build backends, program, websites.

All our target is for the actual code, the codes, the scripts, uh, all these codes. The website is what we want. But our colleagues in, our colleagues in my company, they just use Cursor to do data visualization and, uh, fi-- uh, ba-batch file processing and, uh, turn a video into an audio.

Yeah, things, things like that. And they don't care about the, the code. They just use it once and they never saw it again. Yeah, things like that. So, you know, after we saw these, um, interesting findings from our long coder, uh, colleagues, we think that maybe we should use a-- maybe we should do the opposite.

We should do the opposite. Yeah, it's like we want to keep-- we want to do the right panel in Cursor and hide the le-left panel, right? Yeah, because they don't care about that. Yeah. And even further, we want to do the right panel in the cloud, which is the learnings from our browser project.

So that's it. That's the end, that's the end of last September, uh, we ended our browser project, and we learned something from the Cursor. So that's the moment that we decided to start Manus. That is, uh, last, uh, October.

Building Manus31:23

Tao Zhang31:23

Yeah. So Manus just started like five months ago, uh, in la- last October. We decided to start this project. There are three key important decisions. One is that we want to build the right panel of Cursor in the cloud.

Yeah. So that's how-- That's why we want to build Man- uh, Manus. Yeah. And, uh, after we decided to build it, uh, then how we build Manus? Actually, it's really simple. It's very simple. As I mentioned before, we think LLM is actually very good, very smart.

But the problem is we don't give this smart person a computer. We only give it pens, pen and paper, and ask it to write down all the ideas inside its brain, but we don't give the computer. That, that is not modern days, right?

Yeah, it's like one hundred years ago. Yeah. So the first thing we do is we give the, uh, we give the Manus a computer. Uh, what, what, what do I mean by this? It's like, uh, we're using an open sourcing, uh, project called E2B.

Uh, if any of you work on this. We use E2B to... Uh, you are smiling. You understand?

Guest32:35

I think they're friends.

Tao Zhang32:36

Okay. Yeah. So we use E2B to, uh, build, uh, to assign each Manus task with a virtual machine on the cloud. And may- I know some of you may ask why, why are you using, uh, E2B? Why are not using just using some, uh, container solution like Docker, things like that?

Yeah. Oh, sorry. Yeah. The decision we made is because we think in the long term, in the long term, uh, to give it a full functional computer is very important. If you are going with the Docker way, there are so many limitations.

Yeah, like in Manus, yeah, if you're, if you are trying Manus right now at the, uh, right top of... there is a, a button we, we, we hide there because we don't want to we don't want the average user to, to use that.

But actually, you can just go direct, uh, into Manus' computer and to see all these functions. Like right now we are giving, uh, Manus access to the virtual machine's terminal and the VS Code browser, and in the future, and not, not very far future, maybe just, uh, three months or six months, uh, later, we will let Manus can use other software on the virtual machine.

And actually, you know, it's really can-- You, you can really imagine what, what will happen. Yeah. And, uh, uh, with E2B's, with E2B's, uh, uh, capability, uh, in the future, actually, we, we are not just, uh, can create a deep lab space virtual machine for Manus.

Yeah, if there are some software that needs to run on Windows, uh, actually we can get you or maybe on Android. Yeah, so things like that. So the first thing is we give it a computer and, uh, give it, uh, some softwares on the computer.

So Manus can control this software on the virtual machine. So each task on Manus has its own computer. That's the first, uh, thing. Uh, the second thing is that we give it data access because, you know, right now the world, uh, uh, uh, not, not every information and every data is on the open internet.

Uh, there are some like internal data, and there are some paid data you can't find on the internet. But for some use cases, if you don't have that, this data, the agent can't get that done. It's like if you hire someone, you hire an intern to your company, but you don't, you don't open your internal system accounts for the intern, the intern can't accomplish his, his own work.

So the second thing we do is like, uh, we, we, we, we buy some APIs ourselves. So we, we prepay, uh, for the user. Uh, so user won't care about all the details like, uh, the access to, uh, some like stocking data.

Yeah, you can just, uh, uh, find the stock data for media, yeah, so for find some real, real price for the... a real-time price. And also the Twitter and LinkedIn, uh, search, searches. Yeah. So if people want to do some research on social network, it can perform kind of these tasks.

That's the second thing. We give the data access. Yeah. And the third thing is that, uh, we give it some training. Uh, that's, that's very interesting because, you know, uh, once you, you hire an intern in your company, I think in the first week there must be so many con- con- conflicts, right?

The interns, uh, he, he or she doesn't know, uh, what's your, what's your, what's your favorite and what, what, what's your, what, what's your like, uh, uh, like, like, uh, what, what you like in, in, in the work output format, things like that.

So you always back and forth, you tell the intern, "Okay, uh, next time if you are delivering, uh, this kind of documents to me, remember to add some sections before blah, blah, blah," things like that. So like in Manus we have a, we call it like a knowledge system, which is like during playing with Manus, you can always teaching Manus What is your own, like, favorite?

Uh, like, uh, in our, in our demo video on our website, uh, Pete just, uh, uh, tell the Manus, "Next time if you are doing the resume screening, just, uh, deliver, uh, deliver a d-d-deliver a spreadsheet. Uh, not a, not a document, just deliver a spreadsheet to me."

And then Manus will remember, next time if you give, give it some task like resume screening, it will always return a spreadsheet to you, but not a document. And also, like if you are, if, if you are using, uh, Manus to do some research and there is someone, uh, you care, you care about a lot, so you just maybe, maybe you can just give a short name like, uh, "Next time I ask you to do some research about Sam," uh, it's, uh, Sam Altman, right?

Things like that can always teach Manus and it can remember this knowledge. Yeah. So that's the third thing, is give it some training. But these three key pillar components of Manus, I think they, they are important, but not the most important.

Yeah. I think what's the most important thing is, uh, is the fundamental concepts, uh, beneath the Manus. It's less structure, more intelligence. Because, you know, uh, uh, Parker and I, we, we spent, we, we spent one week's time in, in Bay Area, and we've already talked with a lot of agent start- agent startups and agent researchers in Bay Area.

Less Structure37:46

Tao Zhang38:13

And, uh, most, uh, people I think right now they are working on some like predefined workflows solution. Yeah. Because, you know, people think, "Okay, we need stable and we need accuracy, so we, we must have some predefined workflows."

But when we build Manus at the very first day, we are targeting at general agent. We want to build products for average people. Yeah, for like, like normal users. So the first decision we made is that we are, we are not making another coding agent like Cursor and Windsurf.

Yeah, like, like, you know, a lot of people want to do. We're not doing that. We are building for average user, for normal users. So we want Manus can do like universal tasks. Yeah. So if we are choosing the predefined workflows way, we have to build a lot of workflows for it, maybe hundreds of them.

Right now you can see from our website and also from our Discord servers, people are just using Manus for very different kind of tasks. Yeah. A-actually, we-- I think we can't prepare enough workflows for that. So at the very heart of Manus, actually we just, uh, keep it at very simple but very, like, sophisticated structure.

Uh, and it just gives it more intelligence. Yeah. So it's kind of like provide more context to the LLM and, uh, not try to control the LLM's thinking. To give the thinking to the LLM, we just keep focusing on the hand work.

Yeah. The, the environment building. Yeah. So that's the story, uh, behind Manus for the past five months. Yeah. Yeah, if anyone of you, of you have questions, uh, we, we can just do a very quick, very quick Q&A.

Yeah. Okay. That, that's, that's all for today.

Quality & Context40:17

Guest 240:17

So I, I think I, I really-- I think the idea of building a general agent is really good. Uh, I have, um, uh, two quick questions. One is, um, how do you, how do you improve? Because you-- once you release the product, you keep quality, right?

Tao Zhang40:33

Yeah.

Guest 240:34

Building a general product, you gotta keep quality. And management report is gonna be a very challenging task.

Tao Zhang40:39

Yeah.

Guest 240:39

How do you handle that? Secondly is a lot of these agents, one of the challenge is the context length, right?

Tao Zhang40:44

Yeah.

Guest 240:45

Building a genet-- a general-purpose engine, you gotta handle so many different tasks, so many different tools, the logs. How do you handle the, that challenge?

Tao Zhang40:52

Okay. Uh, for the first question, uh, right now, uh, what we are, we are working at is, um, we, we will face some... There are, there are people, they will report some issues to us like, "Okay, in this, in this task, I think Manus won't, won't work with that."

Uh, but not, you know, not every feedback we got we, we, we would have solved that because, uh, some tasks right now users give Manus are just, uh, too hard. It's like, uh, teach me how to earn one million in one day, you know.

For that kind of request, actually we as Manus can't, can't handle that. But for some tasks, we say, "Okay, this task, uh, actually is really inspiring. If we solve that, many other people will benefit too, and we will, we will invest some time."

And, uh, right now we find the solution comes out from two ways. One is the tool, because you know, right now Manus have like twenty more tools, uh, which is like maybe writing a file, editing a file and brow-browser and URL, things like that.

And, uh, uh, in the past three weeks, we found that there may be some missing tools we are not ready for Manus, and we are building it. Like the first thing, and it's already online two days ago, it's read the image.

You know, two, two days ago, Manus just, uh, uh, Manus can only view the whole webpage. But for some images it downloaded from the internet into the local file system, we don't provide a way to actually look at the image.

And right now we provide this function, and the, the capability of the whole system just go up. So one thing is that in the future we're gonna provide more tools for the agent. And the second one is I just mentioned before, uh, we prepare some data APIs for Manus.

Right now it's very limited. We, we only have like sixty APIs. But, uh, right now we are looking at what users are building with Manus and, uh, we are, we are deciding which APIs we're gonna inter-integrate, integrate into Manus.

So that's two ways, like, uh, tools and, uh, APIs. Yeah, that's, that's the two way. And also, you know, the, the foundation part is like we are always expecting, uh, the foundation model company, they will release more powerful, uh, foundation models, and that's another way, you know.

Yeah. And for the second question, actually you are asking about the context length, right? Uh, you mean how can we handle the long context, right?

Guest 243:21

Yeah. And even, even when you, when you use the so many tools, you're gonna get a lot of feedback, a lot-

Tao Zhang43:26

Yeah, yeah.

Guest 243:26

-of tools. It's gonna become very long unless you give them even more tools. So how does Manus manage to handle all those context?

Tao Zhang43:32

Oh, yeah, yeah. Sure. Uh, one thing is that, you know, this industry just is evolving very fast. So like one year ago, you can't imagine right now we can have this long context, uh, lengths in the foundation model.

Yeah. Uh, so actually it, it is solved by the foundation model part. But even with their, their, their context lengths right now, uh, there are some times you will exceed the context length, right? So we have to do the split for the context, uh, for whole context.

Yeah, but there are s-some philosophy behind how to do the split. It's like, uh, you, you, you cut them all. You just, uh, s- uh, do like phrases, select some of them to be split. And, uh, that, that is some, you know, some research paper here.

Yeah, but we are just, uh, doing the split things. Yeah. Okay. So any other have questions? Okay.

Guest 344:27

Yeah. It looks, uh... First of all, great presentation. Thank you for, uh, the, the synopsis. So y-your costs seem to be mostly on the tokens, right? At this point.

No Own Model44:27

Tao Zhang44:38

Yeah.

Guest 344:39

Do you have plans of building your own, uh, foundational model? Because then you almost will become an open AI, right? What is your plans in building your own generative-

Tao Zhang44:48

Oh, you mean about the pricing plan, right?

Guest 344:51

No. The, the right now, um, you're advertising more like your cost is more on the tokens-

Tao Zhang44:55

Uh.

Guest 344:55

-whatever you're spending on the tokens-

Tao Zhang44:57

Yeah.

Guest 344:57

-for this testing. So do you have plans of building your own foundational model in the future?

Tao Zhang45:03

Oh, uh, actually, we, we don't, we don't have those. We, we don't have that plans because, you know, if you are stepping into the model training area, that will cost a lot more. And we are just a very, very tight startup.

If we are going to the model training part, I think our funding can't support us. And, uh, actually we think in the... Maybe not the very long term, yeah, just, uh, you know, short term, uh, the agen- the agentic capabilities of models will become commoditized.

So it's like maybe you can find so many models in the, in the market that you can use right now. They can support this agentic usage. But, but, but, you know, for, for, for, for right now, there is only as large, like their model can handle, uh, handle the-these, these tasks.

Yeah. So we don't have those plans. That's very expensive. Okay. Uh, yeah. Go ahead.

Guest45:55

I like using Manus for, like, data crawling, uh, things. So like in terms of data APIs, just curious in terms of, you know, a lot of data is behind paywalls or logins. Like, well, I imagine you, you're gonna have to do some partnerships to get access behind login, paywall-

Data Access45:55

Tao Zhang46:12

Oh.

Guest46:12

-data access. Curious, like your thoughts.

Tao Zhang46:15

Yeah. Good question. Um, we are, we are thinking about this question at the very start of Manus because, you know, the, uh, the first pro-- in the first pro-prototype, we found out that, that in the agentic future, uh, the biggest block right now is Cloudflare.

Because if you, if any, anyone of you today, uh, you are, you, you are building agents, I think you must encounter the problem that Cloudflare will block your agents, uh, in any means. Yeah. So we are, we are thinking about that, uh, at the very first of, of, of Manus.

And right now we come up with some solutions to your questions. Um, uh, and, uh, I think the, the solutions are, are in different stage. Yeah. Uh, the s- the solution we are using right now is like, uh, because we are, we are, we are, we are small at this, this stage.

If we go to talk to Cloudflare and say, "Okay, can you, uh, do some partnership?" They, they, they may not answer. Yeah. So right now maybe we can just, uh, uh, have some workaround, maybe using some like residential IPs.

Yeah, things like that for, for this purpose. But we truly have some... We truly have con-confidence, and, uh, we believe that it can, it can be true, which is in the future, it must be an agentic future. It's like everyone will use agents.

So in that future, I think there must be a way for agent company like us to cooperate with these cloud security, uh, providers like Cloudflare. Yeah. To tell them, "Okay, this is not a spam. This is not a bot.

Uh, this is the real usage from real user. It's just performed by agents." So maybe it has, yeah, things like that. As that is one solution. And also we have other solutions is like, uh, uh, maybe we can pay the paywall, right?

There are so many new newspaper websites they, they need you to pay. Actually, we can, we can pay, uh, for our users for these paywalls, uh, because, you know, we, we're gonna release our pricing plan very soon. And, uh, it is a consum-- is a consumption-based plan.

So, uh, you, you already pay for the agents, so we can just, uh, pay the fee for you, uh, by ourselves. So that's, that is another way. Yeah. So yeah. So three solutions. Yeah. Okay. Any other questions?

Okay. So if there's no questions, Parker, we can give the decks part.

Guest 248:56

Does anybody have anything that-