LALatent SpaceDec 31, 2025· 34:45

[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv

AlphaXiv founders Raj and Rayhahn describe building the go-to platform for navigating AI research by obsessing over UI and getting authors like the Laura team to comment directly on papers, then evolving from comments to feeds to benchmarks to AI assistants. They detail how they beat Hugging Face Papers by enabling direct paragraph-level commenting and early traction with high-signal authors, and how they use social signals (views, comments, Twitter) to filter the archive firehose. They discuss the broken academic review process—20% AI-generated reviews at ICLR, quality collapsing under exponential submission growth—and argue papers are becoming less important than implementations, Docker containers, and real-world usability. Their vision: making it trivial to spin up any trending paper's code directly in browser, focusing on the top 0.1% of implementable papers that applied researchers actually care about.

  1. 0:00Origin Story
  2. 2:55Beating HF
  3. 4:50PDF Parsing
  4. 6:00Company Progress
  5. 8:54Beyond Papers
  6. 10:38Paper Picks
  7. 13:13AI for Science
  8. 18:41Agentic RL
  9. 22:22Review Woes
  10. 31:20Power Law
  11. 34:21Outro

Powered by PodHood

Transcript

Origin Story0:00

Rayhahn0:03

Light in space, light 2025. Break ups. Light in space, light 20-

Host0:13

Hi, guys. Welcome to Lane Space.

Raj0:15

Thanks for having us.

Host0:16

You guys are the founders of AlphaXiv, which I've confirmed is the official pronunciation. What's, uh, each of your's, uh, sort of origin story of coming together and starting this?

Raj0:27

Sure, yeah. So I'm Raj. Um, I, I guess all of us were, were at Stanford before. I was roommates with Rayhahn actually, um, and, and still am for last, like, five years. And, uh, yes, um, yeah, we-- I w- started off doing, you know, AI research like a lot of others, uh, at Stanford and, um, was in Dorsa's lab, also working with Percy a little bit and was kind of working at the intersection of robotics and NLP.

And I think, um, honestly, like, the, the real origin story is one day Rayhahn and I, we were in, like, our operating systems class. It was, like, really late at, like, one a.m. Rayhahn's like: "Hey, do you wanna see my, like, final web dev class project?"

And I'm like: "Hey, what makes you think I wanna see that right now ?" We're trying to get this, like, p-set in and go home. And, uh, he was like: "No, you should look at it. It's pretty cool."

And, and I was like: "Okay, sure. Show it." And it was basically, like, uh, like, really jank, like, View Comment button next to, like, every paragraph of, like, an archive paper and the idea is like, hey, you could, like, comment on papers.

And I was like: "Hey, this is, like, kinda cool. We should, like, just, like, work on it." It's kinda strange that there aren't comments on archive papers given that how, how many people are reading, and, like, AI papers is, are growing exponentially.

Host1:41

Yeah.

Raj1:41

So we kinda just, like, started working on it for fun. We showed it to, like, our lab, and they're like, "This is pretty cool." Like, they'd use it if it, if it takes off. But I think, um, it really became serious when, like, it went viral.

Like, one of our friends, um, who was at FAIR at the time posted about it on LinkedIn, went viral, and yeah, from there we were like-- we started working with Sebastian, who is our advisor. Um, and yeah, um, it's been pretty cool.

Host2:04

Yeah.

Rayhahn2:05

I, how you-

Host2:06

Or-

Rayhahn2:06

From line. Um, yeah, my name's Rayhahn. I, before working on AlphaXiv as an undergrad was doing, uh, research, uh, in robotics and RL. A lot of this was still before GPT, so the thinking with, as an undergrad is, as, as a meager undergrad, your PhD mentor gives you papers, and you struggle to understand them.

And so the thinking was that there should be some Stack Overflow analogy for, for archive papers, that surely there are other undergrads out here that have these questions. Um, and obviously since then it's taken many forms, but that was kind of the initial, initial iteration.

I mean, like Raj and Rayhahn, I also did research in undergrad. Um, it was sparse deep learning research with Stanford Arlo's group. And I think, kinda like what Rayhahn was, where it's like we all had this shared experience as undergrad researchers, and it was like, it seems very simple, you know, just, uh, the simple premise of comment section for archive papers.

Host2:55

Yeah.

Beating HF2:55

Rayhahn2:55

Um, and it was just sort of a project that we worked on, you know, 20, 30 hours a week whenever we had, you know, free time. Um, and obviously it's just sort of, sort of launched, um, from that.

Um, and it's been a really exciting journey.

Host3:08

I guess the other-- the, the reason that when I first saw AlphaXiv, I was-- I didn't necessarily think about it as new was because I knew that Hugging Face had also launched, like, a, a paper discussion thing, and like, well, so, you know, Hugging Face is huge.

Raj3:21

Yeah.

Host3:22

Why did you seem to win or break out versus them?

Raj3:26

Yeah.

Rayhahn3:27

Yeah, I think there were a, a few things. Um, even when we were just comments, I think we-- uh, I think the core of it was we were our product's ICP. Like, we knew Hugging Face Papers existed, um, and it felt like some of those problems were unsolved.

So something as simple as being able to comment directly on the paper, the, the user interface. I think when we started, we were very deliberate about getting authors, like Laura-- from Laura, DPO, Llama 3, just-- and actually having really useful, cool exchanges, and then those would get shared, and those-

Host3:54

Did, did you, like, reach out to them and say, "Hey, please do it here?"

Rayhahn3:57

Yeah. We would use our existing kind of like people from, people we knew from research and-

Host4:01

That's rough. Yeah.

Rayhahn4:03

Yeah. Yeah, but, but from there it would grow, like, re-really, really quickly. Um, and I think f-from comments, there's been a really natural progression that we've, we've taken that hasn't been with other kind of paper platforms. So when commenting took off, we were like, "Okay, we have a good idea of what papers people are reading," right?

So if you're trying to discover papers, on one end of the spectrum is, like, the archive sort by new, right? So you have like a thousand papers every day sort by new. And on the other end of the spectrum is, like, Twitter.

So you're like, "Okay, actually this is a good chance to put together some higher signal feed of papers based on the interaction." So commenting became feed of papers, and then from there it's like, oh, a lot of the questions people are asking can be probably answered by AI.

Gemini came out with a one million context thing. Okay, let's throw that on here. So we've kind of progressed and built a more cohesive experience just growing our user base, um, and, and, and pulling that thread.

Host4:50

Just on the Gemini note-

PDF Parsing4:50

Rayhahn4:51

Yeah

Host4:51

... do you parse, uh, PDF to text and then context, or just throw the images in-

Rayhahn4:56

So, so it de-depends on the model.

Host4:58

Yeah, yeah.

Rayhahn4:58

Some of these mo-- like we support a bunch of different models. Some are just text-based. Um, we, with models like Claude, we try to pass in the relevant diagrams as well, so.

Host5:06

What, what's the best PDF parsing model?

Rayhahn5:08

Ooh, that's a good question. In the last month, there's been, like, a flurry of different OCR models that have, that have come out starting with-

Host5:13

Oh, yeah, DeepSeek. How's that?

Rayhahn5:14

So I think in terms of cost and accuracy, DeepSeek is pretty good. Like, if you, if you host it on your own A100s and just, like, s- you, you batch things properly, probably DeepSeek is best bang for your buck.

Um, there are services like Mis- like Mistral that have their APIs. Those are probably-

Host5:28

Mistral OCR. Omo OCR.

Rayhahn5:30

Yes. And those might be a little bit more expensive if, if you're using their API offering, but I think, like, DeepSeek OCR is very, very good.

Host5:36

Any Moon Dream?

Rayhahn5:37

Oof. Uh, I-- probably kind of later, I'd say. Yeah.

Host5:40

Yeah.

Rayhahn5:40

But the, the OCR example is really funny. Like, it's one of the cool things we see on AlphaXiv, which is on day one, you know, someone will release an OCR model, and then the next two, three days, like, four other people will, who've been probably working on this for the last many months will go put out their own OCR model.

Host5:56

Yeah.

Rayhahn5:56

And so it's kind of cool to see the, the, the value that AlphaXiv brings there. Yeah.

Host6:00

Yeah, cool. Um, so w-what's the progress of the, the company? You guys decided it was a success site project.

Company Progress6:00

Rayhahn6:06

Yeah.

Host6:06

When did you decide to I guess take it-

Rayhahn6:09

Yeah

Host6:09

... to the next step.

Rayhahn6:10

Yeah.

Host6:10

And also where are you at now?

Rayhahn6:11

Yeah, for sure. So I think for the first year or two of working on it was a project, and what really forced that transition for us was we were looking at, you know, papers with code, weights and biases, Hugging Face.

And we're like, "Okay, we have a user base that really loves what we're doing." And obviously research extends so much more beyond papers, right? Like, we wanna do things with other artifacts of research. One of the things we're, we're gonna be working on is making it really easy for people to play around with, like, implementations of papers directly on the site.

Host6:37

Are you, are you, uh, fully, full replacement for papers with code already, or?

Rayhahn6:42

I would say so, yeah, 'cause we have like benchmarks and, uh, a lot of it, like the state-of-the-art page and like-

Host6:46

'Cause I know a lot of people miss it

Rayhahn6:47

... it was taken down.

Host6:47

Yeah.

Rayhahn6:48

It was taken down. So, uh, we've added a- like the feed, the benchmarks, whatever. But I think because-- I think the thing that drove us to be a company was there, there are so many different artifacts of research beyond papers and, you know, it, it just kind of made sense to go from there.

Host7:01

Yeah. I don't know if any of the, the others wanna chime in.

Guest7:03

Yeah, I think like, uh, like kind of the analogy of like Arxiv existed, and we came in and like built like an intelligent layer over Arxiv. You know, much nicer interface with tools to quickly understand core ideas on papers, comment.

Like, if we can do the same thing for like benchmarks, right? You mentioned like papers with code used to be good. It was still a lot of like community-driven. Like people would upload their own benchmarks or like, you know, s- um, implementations or whatnot.

Like now with LLMs it's like really easy. Like we, we can just parse through like using OCR, like from charts and tables, like what are the leading like models and papers, uh, ideas on benchmarks and just bring them all in one place and, uh...

And kind of one way to think of it is like just help people make sense of like the fire hose of AI research, and that's just not papers, like benchmarks, you know, models, so yeah, implementation.

Host7:53

Yeah. I mean, the, the thing I really want, uh, which, which, uh, is easy to do now is, um, just like put out a monitor of like when this topic comes up in the ML-

Guest8:01

Yeah

Host8:02

... uh, so like a custom feed of stuff.

Guest8:03

Yeah, that's also where we wanna put a lot of work into our assistant is like kind of have the best research assistant that, you know, knows your interests and can like give you rel- most relevant and up-to-date-

Rayhahn8:14

Yeah

Guest8:14

... like research.

Host8:15

So it's really like a personalization and ranking system, a recommendation system of papers. Anything else that is, is underappreciated? Like do you see yourself as a social network? A Rexis?

Rayhahn8:26

Yeah.

Host8:27

Goodreads?

Rayhahn8:28

No, um-

Host8:28

Genius.com?

Rayhahn8:29

Yeah.

Host8:29

I don't know.

Rayhahn8:30

We've gone through multiple iterations, and I think where we kind of provide the most value is as, as concretely like not a social network, but a kind of a tool over research. And maybe there's some s- some signal that's from, for the-- some type of social signal.

But concretely, people use us as a tool to understand research. And I think one thing that's underappreciated, uh, is that papers are like the tip of the iceberg of research, and I think we're very well positioned to he- help people beyond just the workflow of reading papers.

So you could imagine, uh, one thing we're really interested in, it's not, uh, publicly released yet, uh, but we wanna start maintaining like Docker images for papers and make it really easy for people to spin up implementations directly from their browser.

Beyond Papers8:54

Rayhahn9:06

Um, and you know, a lot of times when people are reading papers, why are they reading papers? Our, our audience is like a lot of applied researchers, so they're figuring out what research is relevant for them. And to get that, you need more than what a paper can tell you.

At the end of the day, papers are great, but they're like a puff piece for the implementation. That's what they are. And so getting closer to what the signal of the research is, is means you need to allow people to play around with the implementation.

So I think one, one thing that's underappreciated is how much potential there is here to do things beyond, uh, papers in this, in this realm of making tools for, for researchers. Um, yeah.

Guest9:37

I think even with the, the current state of the site, I would say that if you want a sort of bird's-eye view on the state of AI research, on the state of AI in general, I think AlphaXiv is the best place to do it.

And if you look at places like Twitter, I mean, it's-- there's a lot of stuff on Twitter, but it's very noisy. If you look at Hugging Face, sure, there's, you know, the model weights, datasets. There's certain pieces here, but it- it's not really a cohesive, holistic experience.

And I think that something that we've done is obviously we started as a comment section, but we sort of, uh, created a platform that anyone ranging from ML researchers to like VCs, if they just wanna stay up-to-date with what's going on in AI, I think AlphaXiv has become that place.

And obviously, like Rayhahn has alluded to, we wanna go deeper. This-- We, we don't want AlphaXiv just to be a place where you, like, you just look at things and you're just surveying things. We want it also to be a place where eventually you are doing your own experimentation, you are actually working with lots of implementations of papers.

Um, and I think that we're just sort of scratching up surface of what's possible here.

Host10:36

So w- uh, we're here at NeurIPS.

Paper Picks10:38

Rayhahn10:38

Yeah.

Host10:38

Uh, you've been here the longest, and you have some like... I asked you to guys to like maybe talk about, like maybe top papers of the year.

Rayhahn10:45

Yeah.

Host10:46

Top papers at NeurIPS, your favorite paper that you can't shut up about.

Rayhahn10:49

Yeah.

Host10:50

Obviously, working at AlphaXiv, you have to love papers.

Rayhahn10:53

Yeah.

Guest10:53

Yeah.

Host10:54

Go for it.

Rayhahn10:54

Like, go ahead. Okay, sure. Um, I think so one category of papers, and I'll start by prefacing, we saw the quote on, on Twitter from, from Ilya that like the age of scaling is done, now it's the age of- I feel like we see some embodiment of that in AlphaXiv where obviously the people posting open research may not always have thousands of GPUs, right?

And so you're forced to come up with like really creative solutions that, uh, to, to problems with limited compute. And, you know, whether these translate to like practical applications or not is maybe a separate thing, but I always find these type of papers very interesting.

So one paper that like for the last month I can't shut up about is, uh, Tiny Recursive Models, TRMs.

Host11:32

Yeah.

Rayhahn11:33

Simplification of-

Host11:33

Related to HRMs?

Rayhahn11:34

Just a simplification-

Host11:35

Yeah

Rayhahn11:35

... of HRMs, which is like biologically inspired. TRM cuts out all the biological inspiration, says, "Hey, let's take like a seven million, uh, parameter transformer model, have it pass its latent vector and, and output itself recursively, um, and use that."

So a seven million parameter model to get like really, really good results, relatively speaking, on really specific but challenging puzzle tasks. So like Sudoku, Arcagi. I believe it gets like forty-five, fifty percent on Arcagi 1 compared to like, it's not state-of-the-art, but you compare that to like DeepSeek R1 or, you know, Gemini 2.5, it's like pretty good.

So I think these types of creative, creative solutions are really fun. They're well-received on arXiv all the time, like when, when these type of papers come about. And so I think just as a, as a nerd for papers, I find that really, really cool.

Another one is like really recent, I believe from like the last week or two, the, uh, evolutionary, uh, strategies are at, at hyperscale. So basically we're gonna cut out gradient descent and use like low-rank perturbations at like million-parameter scale and get like really, really good results.

Um-

Host12:37

So, so use search rather than gradient descent.

Rayhahn12:41

Yeah, some random perturbations. You can't do this over a million parameters at once, right? That doesn't make-- But you can make a low-rank perturbation, sets of low-rank perturbations, and then apply that and so it's evolutionary, right?

Host12:53

Mm-hmm. Yeah, yeah.

Rayhahn12:53

And you get like pretty good results compared to like, I believe like base compared to like GRPO or o- on some tasks and I, I thought that was... Also it was trending on arXiv and I was like, "This is, uh-"

Host13:04

Mm.

Rayhahn13:04

"I love to see works like this." Yeah. And again, whether it scales to anything, I, I don't know, but it's, it's cool to see, uh, this kind of work.

Host13:12

Love it.

Rayhahn13:12

Yeah.

Host13:13

Great recs. Uh, let's, let's keep going, Ram.

AI for Science13:13

Raj13:14

Yeah. Um, I think one area that's also been really exciting, especially among our user base and I'm seeing here at NeurIPS as well is like AI for science. Like-

Host13:23

Mm-hmm

Raj13:24

... kind of, you know-

Host13:25

We're starting a s- a dedicated AI for science pod. It's my, my second podcast.

Raj13:29

That's awesome.

Host13:30

Yeah.

Raj13:30

Super exciting. Yeah. We, we kind of have like a big community as well, and actually we're doing an event on Saturday. Um-

Host13:35

Oh, shit.

Raj13:36

Yeah. Um-

Host13:37

Well, I'll, I'll, I'll bring my new host. Yeah.

Raj13:38

Awesome. Yeah, that'd be exciting. So we have like one of the co-founders of Lila, um, Jane Zou from Stanford, um, Jeff Clune as well, and I think, yeah, like I, I, I was first really excited. I saw this paper, it was earlier this year, called Agent Laboratory, um, from Sam Schmidhuber from DeepMind, and the idea is like, hey, like what if we basically automate the scientific process itself, you know?

Like basically it would take like have different LLM agents for like different parts of the scientific method, like one for like literature review, one for like running experiments, and one for like analyzing those results and writing a paper.

Um, a- and, and it was pretty impressive. I mean, it's still like I think gonna take some time, like a lot of like human, like in the loop still like makes a lot of impact, but like they were able to show that, you know, like with like a few dollars from like idea to like final report for, for certain topics.

So it was like pretty interesting and I mean, we're seeing a lot of companies too kind of forming in this area like, um, Axiom recently, um, kind of-

Host14:38

The math.

Raj14:38

Yeah.

Host14:39

Yeah.

Raj14:39

Yeah. Like it's, it's I think like a really interesting space so.

Host14:42

I have a problem with it, maybe, maybe others do, maybe others don't, of just lumping everything into science.

Raj14:48

Yeah.

Host14:49

So-

Raj14:49

Yeah, it's definitely like-

Rayhahn14:50

It's a pretty-

Raj14:50

It's like broad, like- There's like, you know, like, like you said, like math and then there's like more like the life sciences, which feels like a little bit more difficult. Um-

Host14:59

They have nothing in common.

Raj15:01

Like, um, like, like automating like ML research seems a little bit easier-

Host15:06

Mm

Raj15:06

... than like kind of my in like intuition than like, you know-

Host15:09

Yeah.

Raj15:09

Building like a new-

Host15:10

AI for ML research-

Raj15:11

Yeah

Host15:11

... can we agree that it doesn't count as A- AI for science? Like, like-

Raj15:15

I think that's a, that's a fair take. Yeah.

Rayhahn15:17

Yeah.

Host15:17

Then everything is science. I don't know.

Rayhahn15:19

You can make an argument ML research is not research. Like-

Raj15:22

More of an art.

Host15:24

Okay. So just generally AI for science, any particular paper or discussion that you've had in the last couple days that like sticks in your head?

Raj15:31

Yeah, I think, I think, uh, yeah, like the Agent Laboratory paper.

Host15:34

Oh, yeah, yeah. Agent Lab.

Raj15:35

Yeah, that.

Host15:35

Okay. Yeah, so my comment on that is every quarter-

Raj15:39

Yeah

Host15:40

... there's somebody who will be like, "Oh yeah, we've done AI scientists."

Raj15:42

Yeah.

Host15:43

And then I don't know if it goes anywhere. I don't know if it's like a proof of concept. Like what is going on.

Rayhahn15:50

It depends on the example. Like going back to the AI for life science, I'll make the distinction. One of-- We had a speaker, he wrote this paper called BioOmni, uh, Ke Xing, and, uh, started as a paper. Uh, right now it's, you know, he's kind of, he got a lot of traction and he's kind of continuing to develop it.

I think that was the first time I was like, okay, conceptualized it. People are using it as a tool and obviously I think things are gonna take time. Part of it is maybe to get traction, you call it like, you know, it's, it's we're automating all of science.

But it seems to be like in his case with BioOmni that there are people at pharmaceutical companies that are, are bio organizations that are using it as a tool. Uh, and to what capacity? I'm not a life science expert, but I think there's probably steps from it being a tool to like automating like a lot of things.

I'm skeptical of like the, the automation of the whole process. I, I like the idea of thinking it more of as a virtual lab. So Jane Zou frames these AI scientists as, as a virtual lab where it's like, okay-

Host16:42

It's cheap to experiment and explore.

Rayhahn16:43

Exactly. So the person is the PI and I think that vision makes a lot of sense. I think to replace-- I don't know how much of it is pitching versus science where someone's like, "We're gonna just completely build a scientist."

'Cause in my, in my way, one of the theses of arXiv, like as we go on there's gonna be more researchers and maybe there's more tools to help them, but the actual like process of doing research is like, like human, human work is gonna be like that is one of the final frontiers of human work is to be able to intuit and, and ideate and figure out what's next, right?

So I see these as all like tools, um, that-

Raj17:13

Yeah. I think, I think, I think just extending what Rayhahn said, I think, yeah, maybe for like, you know, pitching or like, uh, what we see in the news, like people are like, "Oh, we're building like AI scientists or like an AI material scientist."

But you don't actually need to do like that whole end-to-end workflow to provide value, right? Like what Rayhahn was men- mentioning about BioOmni, like I know they're like providing a lot of value with like their tools for like, oh, we'll like sort through the lit- literature for you, right?

And like find out what's, what's relevant and then like help you get some experiments set out. You know, that's like you're not automating the whole workflow, but like it's, it saves people a lot of time where like researchers, they're biotech, you know, applied people.

So I think, yeah, like I'm also aligned with you that a little skeptical about automating that whole end-to-end workflow, but I think there's value intermediate parts.

Host18:01

Yeah. Well, skeptical that it is actually put in practice rather than just done for a paper for marketing purposes.

Rayhahn18:07

Yeah.

Host18:08

But yeah, we can, we can move on to other papers.

Guest18:11

Just touching, touching on that one, one last thing.

Host18:13

Yeah.

Guest18:13

Just like a lot of the science that is done is not sound very sexy. It's a lot of like ablations. It's a lot of very sort of like roadwork that's manual and- Obviously, it's like it's not sort of the-- doesn't require the ingenuity or sort of the, the human brilliance, but I think that agents are very much capable of doing a lot of the things that are required for like a lot of scientific processes.

And so I think there's a tremendous value there. Yeah, it's, it's not quite what you're talking about, like a mentat scientist, but I think there, there is something there for sure.

Host18:41

Your paper?

Agentic RL18:41

Guest18:42

Oh, yeah. I think, I mean, for me, I think something that has emerged in the last year or so has sort of been building agents with RL, um, for, you know, long horizon, you know, complex tasks. And I guess one specific paper that, you know, I found interesting was like Agent R1.

And so they basically-

Host18:57

Agent R1?

Guest18:58

Agent R1.

Host18:59

I don't-- I think I missed that.

Guest19:00

Um, I mean, but they, they built an an- an end-to-end sort of RL framework for training agents on s- sort of like these complex reasoning tasks. And the, the specific reasoning task that they chose was sort of like multi-hop, uh, question answering.

And so you have to train an agent to be able to interact with sort of a tool environment and, and sort of asks, uh, and respond to complex queries. And so, you know, when you apply that framework on top of, you know, a base model like Qwen 2.5, you suddenly now have an agent that is, you know, very, very good at these sort of these more complex tasks.

And I think the reason why this is interesting to me is because you see a lot of parallels in this paper to what's actually happening in like industry. You know, like one, you know, one example of this is just like Cursor, their composure model.

There's a lot of rumors online that, that this was built on top of, you know, Qwen 2.5. People saw like Chinese reasoning traces and whatnot. But- I mean, we don't know exactly how they built it, but sort of these, these techniques, you see them-

Host19:54

I see them.

Guest19:55

Yeah, I mean, yeah. I mean, th- it is a very popular technique to like, you know, build on top of Qwen 2.5. But then there's the whole crop of like, you know, RL as a service, fine-tuning as a service, these sorts of things to help, you know, companies build in-house agents.

And so a lot of what we see, you know, in, in enterprise and industry these days is, uh, a reflection of sort of a lot of the research that's done with, you know, AI agents, um, and RL. So I think it's definitely very exciting.

Rayhahn20:18

And one, one thing to add there is like, so there, there's frontier labs aside, there's these c- startups that raise like collectively hundreds of millions of dollars that are all like fine-tuning Qwen, like specifically Qwen. Because like Qwen, if you like a- talk to some of these folks, like Qwen was specifically designed to be like very sample efficient when it comes to like RL fine-tuning.

Like they designed this w-when building the model, and it's so cool to see all these different startups and it's like goes right back to Qwen and like the open source landscape. Um, and then, yeah, you see papers like Agent R1 that-

Host20:47

Was, was that DeepSeek? Literally DeepSeek R1 or that's-

Rayhahn20:51

So I, I don't know why they threw an R1 in the name, but they were fine-tuning-

Host20:54

That's the Qwen model, right? Okay.

Rayhahn20:55

Fine-tune. I know.

Host20:56

I was like, what? Why is it R1?

Rayhahn20:57

Yeah, I don't know. Um.

Host20:59

Uh, Justin here, uh, from, from the Qwen team. Um.

Rayhahn21:01

Okay.

Host21:02

Yeah, so he's roaming around and, uh, I, I think the-- this is very much Qwen's year.

Rayhahn21:07

Yeah.

Host21:07

I think. Uh, in, in fact, comparatively, I think DeepSeek actually started the year super high-

Rayhahn21:12

Yeah

Host21:13

... and then fell off a little bit, which is kind of interesting.

Rayhahn21:15

Yeah. I was just talking with o-one of the, uh, folks here about presence of DeepSeek as far as like what we see on arXiv. They still pull a lot of weight in terms of when they put out work, people listen.

Um, even if-

Guest21:27

It's great work

Rayhahn21:28

... it's, it's great work. So-

Guest21:29

Great papers and work

Rayhahn21:29

... an example is with OCR, right? They put out DeepSeek OCR, and then like the next day, like Olmo OCR 2 comes out. The, the startup called Chandra puts out an OCR, uh, model. So-

Host21:39

Yeah, but if I, if I'm not mistaken, their numbers were worse than DeepSeek still.

Rayhahn21:43

Yes. So but, but I think it still goes back to the point of like when DeepSeek publishes, other people are very much still intact. So whether or not they-- I don't know. It's, it's short term-

Host21:51

I don't know

Rayhahn21:51

... if it's fallen off or not, but, uh.

Host21:52

Like difference of one day, it's not, that's not motivated by the DeepSeek. I don't know.

Rayhahn21:57

No, right. But I think they feel press-- it's, it's like the parallel between research and like product releases almost is like-

Host22:02

Okay, yeah

Rayhahn22:03

... you kind of see that here.

Host22:04

Yeah. I mean, look, a lot of interesting papers and, and, uh, coverage in NeurIPS. We j- we just covered, uh, best paper for Deep RL. I think another like what I'm trying to make kind of exploring is like just the health of the academic paper ecosystem, which obviously you guys care a lot about.

My sense from talking with area chairs and program chairs is that there's a lot of review process is not keeping up. Therefore, the conference quality proce-- uh, submissions and quality is just going down. ICLR had a big scandal recently with leaking names and like twenty percent of their reviews are AI.

Review Woes22:22

Rayhahn22:41

Yeah.

Host22:41

What's your take on just like all this?

Rayhahn22:43

It's, uh, it's a shit show for sure. So there's-- I think on multiple levels. Um, one fun project we've wanted to do on arXiv, I'm skeptical on if w-we-- this would be good for us or not. But first of all, we've always been hes-he-hesitant to put out tools that allow people to write papers with AI.

Host23:00

Mm-hmm.

Rayhahn23:01

We help people understand papers with AI. We feel like something feels wrong about like, okay, people are gonna write papers, which they're probably already doing. But one thing we want to put out is which papers that are trending are actually like AI-generated, and you'll probably see at a very high scale.

A lot of these-- there's a lot of pressure to put out lots of work and do it really fast. And, uh, there's a lot of pressure to write papers quickly, and, uh, it's a very much AI in, AI out thing that we're going into, and so it's definitely problematic.

And yeah, I, I guess there's pressure to publish, and then there's the AI tools that help people write papers and, uh, the quality of the paper itself I feel, uh, has probably taken a hit, but-

Raj23:36

Yeah, I think it's also interesting, um, like I know some people are also working on like AI reviewers too, which is also interesting. And I mean, I have some concerns, but I think overall if done right, it could be good for, for authors actually if like authors get access to this and they're able to like just run their papers through these reviewers before and they can see like, okay, like what's maybe unclear in my paper and like what can I do to make it better?

Host24:04

Ah.

Raj24:05

And, and maybe that just like now the reviewer doesn't need to spend as much time just trying to like understand and parse through and can just like actually point out like or, or do like a literature review and like understand.

Host24:16

Yeah, it's like a linter like quality check.

Raj24:19

Yeah. Yeah. If it's not done right, then maybe people will like find a way to like hack around their, you know, AI reviewer. Um-

Rayhahn24:25

Yeah

Raj24:25

... I know Andrew Ng recently put out like the Stanford Agentic Reviewer, which has been getting some traction. But, um-

Rayhahn24:32

Oh, I wasn't aware of that. Are you guys adopting it?

Raj24:34

Uh, we haven't really thought too much about it yet. Yeah.

Rayhahn24:38

I think, like, one of the tools we'll want to incorporate is some type of reviewer just for the assistance of-- for an author before they publish. Like, "Hey, let me-

Raj24:45

Yeah, yeah

Rayhahn24:45

... look through." We can tell you how much is obviously AI-generated, what parts are... like, what works are similar, and kind of do a breakdown. I, I think the linter analogy is good. Uh, uh, like, I think assessing novelty is probably tricky, but as a first step, I think it can, like, weed out like either low-quality work or-

Raj25:03

Yeah

Rayhahn25:03

... I think that's probably promising.

Raj25:04

Yeah, I, I, I was wondering just specifically on novelty and lit review, how important is search and your, presumably you, you have a index and you-

Rayhahn25:15

Exactly

Raj25:15

... you do rag over the index-

Rayhahn25:16

Yeah

Raj25:16

... and all that.

Rayhahn25:17

Yeah.

Raj25:17

Um, are models good at it?

Rayhahn25:19

So I, I guess it depends on the type of queries. I, I guess when it, when it comes to novelty, it's concerned like arXiv has had this problem, like even early on, where people would use it for like, I forget the exact term, but people would post a paper even when it's not fully developed, and then you kind of claim-

Raj25:36

It's for preprints.

Rayhahn25:37

Right. Right. For preprints, right. So I think in terms of finding relevant literature, it's, it's good. One of the things we kind of, where our assistant defers is we'll use social signal as well. So maybe someone has some paper that wasn't, that didn't get a lot of traction and, um, I think just building semantic search isn't super helpful.

Um, but yeah.

Raj25:56

So like Twitter or other signals?

Rayhahn25:58

It's a combination. So like-

Raj26:00

You don't, you don't, you don't reveal the for- the-

Rayhahn26:01

So, so we'll use like view count, semantic signal, whatever. Basically if you-- I s- I was kind of rambling there, but if you tr- just try to build a semantic assistant, things get noisy fast. So like we have an index over like-

Raj26:11

These people throw in buzzwords.

Rayhahn26:13

Yeah, exactly.

Raj26:13

Of course.

Rayhahn26:14

Right. Right. So, um, we have like three million papers in arXiv. If you just try to use semantic and given some query like what are papers that do agentic reasoning, you are not gonna find the relevant work. So what we do is like we know what papers people are reading, and that's a factor that weigh- weights in when we, when we pull up relevant results.

This is something we've talked about with other, like other people have built like semantic searcher, like semantic search tools on top of arXiv. Inevitably it sounds good, and then it just ends up like not returning relevant info. Again, the, the reason I was kind of rambling before is we build the assistant more from the angle of literature review rather than assessing novelty.

So, so-- And those are maybe two different things. If you're purely trying to assess novelty, maybe the social signal doesn't matter. In terms of building like a really good literature review assistant, you need some weightage on like, at the end of day these are three million papers, a lot of them are low quality, so you, you need some signal.

Um, I have my doubts on whether just Twitter is the, the, the best form of social signal there, but we have our own in-house.

Raj27:11

I'll also shout out Emergent Minds-

Rayhahn27:13

Yeah

Raj27:13

... uh, which has turned YouTube into a signal, which are very- Yeah, they, they also-- We, we, we like the guy Matt. He like, uh- Yeah. He's part of our LeanSpace Discord. Yeah.

Rayhahn27:21

Oh, nice. Yeah, he's so-- he also has like the AlphaGuy like stats on.

Raj27:26

So you can sort-

Rayhahn27:27

Yeah, Alpha 各项 search.

Raj27:28

Ah.

Rayhahn27:28

Yeah.

Raj27:28

I see, I see that.

Rayhahn27:29

We have like a lightweight collaboration.

Raj27:31

But the, the YouTube thing is his thing, not, not you.

Rayhahn27:34

Yeah, yeah, yeah.

Raj27:35

I, I find the YouTube thing very, very useful.

Rayhahn27:36

Yeah. Kind of extending off that, like I think, uh, you know, nowadays research, it's, it's not just the paper, right? It's like a whole package of like more and more people are opening the code base, which is great and, and we hope to continue seeing that.

And so like there's the paper, GitHub if it's available. People have like a tweet thread. There's like a website. If it's like robotics for instance, there's like video demos. Like there's so much more than just the paper. So like, you know, one day if we can index all of that into our assistant, kind of make it as good as possible for like finding all of this relevant research information.

You, you already mentioned, you know, we're doing papers with codes for like benchmarks as well. Like I think that will be really powerful.

Raj28:15

Yeah. All, all I will say is a lot of people are cheating in the GitHub thing where they only put, put like a README or something.

Rayhahn28:20

It's-- We, we see those, and we make sure to not-

Raj28:22

Okay. Yeah, yeah, yeah.

Rayhahn28:23

One thing to build on that is in the next five, ten years, like, uh, the paper as an artifact is gonna be less and less useful. Like I, I think like research will move away... And we're feeling that here, right?

Like at the end of the-- I was saying this earlier about what a paper is. It's like describing what the new implementation is, what the numbers were, but that's an implementation. Like how much does the paper tell you there?

But I think having really organized, useful implementations is gonna be the future of research. That, you know, a paper is just, a PDF is such an antiquated way of sharing information. Um-

Raj28:51

It's Lindy.

Rayhahn28:52

Right. And that's why we exist, and we kind of make that more frictionless.

Raj28:56

It, it also like totally aligns with what we see in our user base today. Like, we have like tons of users who are like not necessarily trained researchers. Like they're not in academia. They're in industry. They're not at like the big labs.

They're like kind of what we call like applied AI researchers or like research engineers and like they're at companies, you know, like for instance like they're obviously big companies, but they're obviously-- there, there's also like companies like Spotify, Expedia, Nintendo, where it's like you don't traditionally associate with like AI research, but like as part of their job now, like they need to keep up to date with the latest research for like some product or feature that they're building, right?

And it's like, to be honest, a lot of them don't really care about papers. They're just like it's part of the job and like for them it's like really hard starting with what paper should I be reading to, okay, like, how do I quickly understand the core idea?

And what they all ultimately want to know is like, will this work for like what I'm trying to build, right? And like getting to, from that like paper to implementation is like a huge gap right now. Starting with like, like you said, yeah, people are cheating on the GitHub.

Like it's not actually useful. Or like even if it's there, like a lot of times, like setting up the code base with the right dependencies is like a huge pain. So we're thinking about, Rayhahn mentioned like doing Docker containers for papers and like-

Rayhahn30:06

Yeah, that was the original idea for Replicate.

Raj30:08

Yeah.

Rayhahn30:08

They pivoted, I guess.

Raj30:10

Yeah.

Rayhahn30:10

And then I think now I would point to Harbor. I don't know if-

Raj30:14

Man

Rayhahn30:14

... from the Terminal Bench folks.

Raj30:16

Yeah.

Rayhahn30:16

That would be an interesting equivalent. Not, not, not paper focused, more RL environment focused, but you could reuse the infra.

Raj30:23

I mean, yeah, this just going back onto sort of the whole like, the whole paper discussion, I think like you just have to look at the graph of like arXiv submissions the last like ten years. It's just exponential.

Rayhahn30:32

And you might need to like the, the CS-

Raj30:34

Exactly. The CS, AI kind of like-

Rayhahn30:36

Thirty thousand a month. It's great.

Guest30:38

Yeah, I mean, like, and you can have things like an, you know, like an AI agent to review this stuff for like peer review. But like, ultimately, I think the, the real-- the end state here needs to be that we just don't care about papers as much.

Um, at the end of the day, like all researchers want is to-

Host30:53

Useful ideas

Guest30:54

... yeah, just open source the repo, open source the weights, and give me a sandbox where I can v-quickly verify your thoughts or verify whatever ideas and whatever contribution that you're making. And the, the easier that you make that loop instead of having to like read the paper and then like do a lot of nonsense to work through their like slop repo and then try to replicate it yourself, faster you close that loop, you, you'll see a lot less of this like slop submission and-

Host31:16

Yeah

Guest31:16

... stuff like that. And I think it'll be a lot healthier for ecosystem.

Host31:20

Okay. Like obviously you guys do what you, you think is best. If I were a friend and advisor to the company, I would be like, "You might be biting off more than you can chew."

Power Law31:20

Guest31:29

Yeah.

Host31:29

This is impossible task.

Rayhahn31:31

Yeah. So one thing about the point about Replicate, which I was just gonna add, which is rather than viewing this as like a reproduction or like a, a, a reproducibility problem of like, "Oh, we need to try to re-re-verify the results on all these papers and make it easy," uh, I think we go back to the point where sort of user base or applied researchers, they don't care about which works have the best numbers in their paper, but they care about like which works are easiest to implement, right?

And so there's some subset of those and like, can we make it easy for those folks? So if there's, if there's an academic that says, "Hey, I want this arbitrary thing implemented," it's like, "Hey, we can't do this."

Um, u-use the fact that there's some power law here in terms of implemented papers that people really want and make that experience really easy. For example, e-even think about... We were talking about the other day with one of these, uh, one of the researchers at NeurIPS who was like, "I haven't actually like played around with Laura."

I was like, like, "What has stopped you from doing that?" It's like, can we make just super simple on really popular works, make it really easy to iterate on that directly on top of-

Host32:25

Yeah

Rayhahn32:25

... the platform and, and whether there's, you know, some obscure work that has g-great numbers, like whether we can reproduce that is not in the s-you know, space of things we care about.

Host32:35

Specifically, if you have anything involving GPUs-

Rayhahn32:38

Yeah

Host32:38

... I would call out Launchables from NVIDIA.

Rayhahn32:40

Yeah.

Host32:40

Uh, they-- that's came through an acquisition of Brev-

Rayhahn32:42

Yeah

Host32:42

... which I used to be a, a investor in.

Rayhahn32:45

Yeah.

Host32:45

Yeah, I mean, there's like domain specific, uh, ways to do these things, but they're not-

Rayhahn32:50

Yeah

Host32:50

... widely socialized.

Guest32:51

I think the last thing Rayhahn briefly touched on is like this power law that we see a lot is like archive has over two point four million papers, but there's a huge power law on like what are the actual like most viewed papers and it's very long tail, right?

And it's like if we can just like kind of focus on like what would be most helpful for the vast majority of people and just make those really easy for people to build off of, I think that's like a good starting point.

We don't need to do this for every paper, you know? And even if initially it's like a little bit of like scaffolding, like I think if you're saving time for this like exponentially growing group of like research engineers, like-

Rayhahn33:27

One super last point here, which is, uh, you know, one thing we're working on is having an agent that can autonomously set up, you know, a, a Docker container for a paper based on the code base. And now let's say on a certain set of papers it doesn't work and it's gonna need handholding, right?

One thing down, down the line is the onus of this we wanna-- kind of wanna put on the author, which is, hey, we rank papers by implementation ease.

Host33:48

Mm.

Rayhahn33:48

So that's what applied researchers care about, right?

Host33:50

Yeah.

Rayhahn33:50

So if the agent can-

Host33:51

You just need to incentivize

Rayhahn33:52

... if the agent-- already we have authors that come to us saying like, "How do I rank higher on the AlphaXiv feed?" Or like, for one, if you don't have a GitHub code base, like you should add one and that'll help you.

But you can imagine now if people are-- our user base are applied researchers, so they care about only implementation ease. We sort it by implementation ease and whether the agent can knock out some hard, uh, you know, some issue setting up paper, well, that's now on the onus of, uh, of, of, of the author there.

And our user base wouldn't be super interested in implementing that in the first place.

Outro34:21

Host34:21

Awesome.

Rayhahn34:22

Yeah.

Host34:22

Very good cause. I happy to promote it and, uh, congrats on, uh, your success so far. We'll see what happens in twenty twenty six.

Rayhahn34:29

Yeah, I appreciate it. Thanks for having me.

Guest34:31

Thanks so much.

Host34:31

Yeah, thank you.