I've been at two startups that have done genuine world first fundamental research.
The first tried to publish novel results for 3 years in tier 1 journals before finally doing a preprint and telling the tier one publishers to jump in a fire.
The second, and ongoing, isn't publishing anything because of my experience with the first.
That and avoiding openAI and Anthropic copying our results and leaving us with nothing to show for six months of work. The papers only come with the pitch deck.
eikenberry 9 minutes ago [-]
Why wouldn't they just skip the journals and publish the papers themselves? Sharing the research is the important part unless these are academics who are trying to get tenure, grants or such.
londons_explore 35 seconds ago [-]
[delayed]
noosphr 5 minutes ago [-]
The article is complaining that research isn't peer reviewed and research is turning into blogs. I'm explaining why that's the best outcome possible and why in a field as hot as large dl models you won't even get that.
The article is vague about the companies in the paper, for some reason.
In the paper, OpenAI is at the top of the chart for cumulative citations. MEGVII, Hugging Face, Waymo, Momenta, Preferred Netowkrs, Anthropic, Owkin, and Databricks, and Aibee follow (in that order). Yes, that is citations, not publications, but they explain that they're trying to use that as a proxy for significance, albeit an imperfect one.
Companies like Google aren't included because they aren't unicorn startups.
larodi 38 minutes ago [-]
Google published a lot. I still wonder this original transformer paper, and the attention, etc. Why would they allow it to let go in the open? Perhaps because was intended for translation first before someone decided to loop it over itself? Or was so obscure to fellow researchers what do they actually publish?
willy_k 59 seconds ago [-]
It took 6 years to get from Attention Is All You Need to a consumer product. Up until ChatGPT’s release, LLMs were obscure and really were just fancier autocomplete. Deepmind was largely a speculative research division until Google decided to play catch up, and even that took a bit of time as they figured out how to proceed in a way that wouldn’t cannabalize search. So in short they didn’t realize the potential. Thats my somewhat naive take.
dominotw 1 hours ago [-]
databricks is a top ai startup?. i thought they did spark hosting or something.
what makes them a top ai startup
tfrancisl 1 hours ago [-]
$160 billion valuation and market capture among some top (non-AI, traditional enterprise) companies, I think.
1 hours ago [-]
egonschiele 2 hours ago [-]
As far as I can tell, the paper never actually mentions the companies who aren't publishing papers. Open AI, Anthropic, and hugging face are all specifically mentioned as companies that do publish papers. Just FYI for anyone else who reads "AI's top startups" and immediately assumes OpenAI and Anthropic.
randomImmigrant 2 hours ago [-]
What the blogificafion of AI research has done is allowed all kinds of claims and terminology related to AI to be introduced and taken up in a manner replicating social media dynamics. And that is simply not healthy. We’re fast reaching a place where any claim can be backed up with a set of numbers from a number of experiments run in some gamified environment or the other, with little concern for if it all adds up to anything.
It’s a vicious loop, because this same junk then goes in to train the next models which help spit out the next set of models AND blogs/papers.
The net effect is not dissimilar to setting termites loose in a library.
fultonn 9 minutes ago [-]
> We’re fast reaching...
We already arrived at that destination a decade ago.
Probably even further back tbh.
There's just a lot more people playing the game now, without the social indoctrination that made it more tolerable in some circles.
It's not that bad, though. September is annoying but you kind of miss the eternal renewal once you're out.
gowld 2 hours ago [-]
How is that different from traditional research publication? It's faster?
randomImmigrant 56 minutes ago [-]
Top line difference is speed. The impact of speed is deeper and gets felt over time.
Traditional research publications are no angels. They gatekeep research, and also allow financial incentives to drive them to publish junk with their stamp on it.
But a flood of papers doesn’t actually mean more knowledge. In bypassing this route entirely, AI has swiftly lost the ability to engage with itself as a field. And the cost of that is only beginning to be felt.
shimman 1 hours ago [-]
No traditional research is mostly done by post-docs and have a phd level of education rather than a tech bro that passed leetcode. That seems like a good bar to have; not too mention the whole peer review thing, hard to really understand anything if you purposely withhold it and tell people to kick rocks.
greazy 60 minutes ago [-]
Plenty of research is performed by non PhDs.
The main difference is reproducibility and peer review process.
janalsncm 55 minutes ago [-]
Two separate axes here: whether the research is in a blog, and whether it’s peer reviewed.
On the first, I mostly don’t care.
On the second, that’s mostly unavoidable if they want to keep their IP.
sndgndgndgndy 57 minutes ago [-]
Science is a method of inquiry, not an institution or group of people with titles.
s1artibartfast 47 minutes ago [-]
Sure, but it is good that people actually adhere to that method if inquiry.
They are not perfect, but institutions like peer review or universities do provide a framework and incentives for quality.
TimCTRL 2 hours ago [-]
Yet none of them would have been here if Google hadn't published "Attention is all you need", the irony.
dumpsterdiver 27 minutes ago [-]
The irony may only just be unfolding - check out this paper from 2015:
Google may consider the standalone frontier-model arms race economically irrational, while still considering frontier-model capability strategically indispensable. Its longer game is probably not to avoid building the biggest models, but to build only enough of them to serve as capability factories—then turn that intelligence into a much larger population of cheap, purpose-built models.
(Human again) If Google knew what they were on to, why wouldn’t they make it their secret weapon from the start? I suspect it’s because they predicted there would be an arms race, and knew how to profit from it. They had a distillation paper published before “Attention is all you need”. In hindsight, is it ironic at all? Or is it obvious?
pixl97 27 minutes ago [-]
Even Google wouldn't be here.
It was kind of like radiation science before WWII, it would be freely published because it wasn't potentially world changing yet. After it became a government interest, even people doing things unrelated to weapons would be much more apt to hold their work close.
arjie 2 hours ago [-]
The authors of that paper are all at AI companies not named Google.
gowld 1 hours ago [-]
They wouldn't be if Google hadn't published "Attention is all you need".
usef- 58 minutes ago [-]
Lab employees do seem to be "knowledge sharing" as they move employers as far as I can tell, so I don't think that's true,
mountainriver 26 minutes ago [-]
Oh that built on a bunch of ideas that weren’t far off. It was a great discovery but like many it’s standing in tall shoulders
2 hours ago [-]
paxys 2 hours ago [-]
I’m not sure why this is so surprising? AI company does not automatically mean research company. The vast majority of new startups popping up over the last few years have commercial motivations, and use models built by someone else. Why are you expecting them to publish scientific papers?
50% of startups contributing to public research is actually a crazy good outcome. That’s far more than I had expected.
HDBaseT 50 minutes ago [-]
When the second biggest player is named 'OpenAI', I would expect some papers to be published.
antonvs 45 minutes ago [-]
OpenAI and Anthropic are identified in the article and study as top publishers of AI research papers.
KoolKat23 2 hours ago [-]
Perhaps I'm imagining it but the entire industry was build on published research, this "AI wave" is at odds with that and seems to be driven by greed (although they'll claim some arms race or something to help themselves sleep at night).
There should be a new ESG (Environmental, Social, and Governance) policy being pushed recognizing the important role this plays. Although ESG and all norms have been set aside in this grim new world it seems.
fultonn 1 hours ago [-]
> Perhaps I'm imagining it
You are not. A disproportionate amount of value in the computing industry was created by <strike>geniuses</strike> decently smart people who worked together and who decided to just tell people how to do things instead of trying to capture the value of being the first person to figure out how to do those things.
This observation pre-dates the current wave of AI hype by a half century or so.
> driven by greed
I can only speak for myself.
For me it's exactly the opposite. If I want to explain how something works, I can just... do that. If I want to share an artifact demonstrating how to solve a particular type of problem, I can just... do that. If I want to mentor/teach, I can just... do that.
Doing those things within the confines of Academia Approved Institutions is exhausting and distracting.
To wit, and the actual point of this post: the term "Publishing Research" in this article doesn't mean "post it on a .html page and share the source code". It means engaging in a very specific and peculiar and extremely political modality of communication.
And it really only makes sense to do that specific and peculiar and political thing you're at a stage in your professional/personal development where you need to play that particular prestige game. (Which there's nothing wrong with, but it is a deeply cargo culted version of the actual scientific process.)
blackqueeriroh 36 minutes ago [-]
Let me help rewrite your first paragraph:
* A disproportionate amount of value in the computing industry was created by g̶e̶n̶i̶u̶s̶e̶s̶ decently smart people who worked together to do things that seemed impossible and who decided to just tell people how to do things instead of trying to capture the value of being the first person to figure out how to do those things.*
fultonn 6 minutes ago [-]
thx for notes; no complaints on my end.
edit: edited.
83642736392 9 minutes ago [-]
> Although ESG and all norms have been set aside
Fortunately
Aurornis 1 hours ago [-]
> Perhaps I'm imagining it but the entire industry was build on public research
There is some excellent publicly-funded research in there, but pivotal papers like Attention Is All You Need and the numerous pivotal OpenAI publications were privately funded.
OpenAI is at the top of the chart in the study.
I think you're bringing some assumptions into this conversation that aren't supported by the evidence.
anon373839 1 hours ago [-]
> OpenAI is at the top of the chart in the study.
A distinction should be made between the old nonprofit OpenAI and the current organization. They don't publish technical research anymore.
KoolKat23 1 hours ago [-]
Sorry I mean published i.e. available to the public.
Aurornis 1 hours ago [-]
I was also referring to published papers available to the public.
Like I said, I think you're bringing some assumptions to this conversation that aren't based on the how the industry came about.
1 hours ago [-]
janalsncm 46 minutes ago [-]
Some companies are sharing their research. Mostly not the American ones
When they do this, are they buying a single copy of say Book A and destroying a single Book A copy? Or are they buying every copy they can find of Book A and destroying all of them?
xyzzy123 27 minutes ago [-]
Hi, shredding the book is to reduce potential legal liability of format shifting - if you shred the book after scanning, the theory is you are not increasing the number of copies. This has come up as a factor in legal rulings. The legality of format shifting is still murky though, and the legality of training on the work are a separate question.
ComputerPerson 2 hours ago [-]
Elon recently (2d) tweeted about making sure they preserve rare books and "scan them the hard way". That tweet launched a cultural wave of opposition regarding the debinding of rare books.
The discussion has been around, but it's flared significantly recently. Not sure if that's what the person you're responding to is specifically inflamed about.
Regardless, old/rare books are certainly being acquired and destroyed.
gowld 1 hours ago [-]
No, he didn't tweeted about "making sure" of anything.
Elon tweeted "I’ve asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning"
which is no sense a credible source for what he actually asked SpaceXAI team to do, or if they are doing it, or what they were doing before yesterday. Elon has an extremely long and thorough track record of lying.
The problem remains either way. Small operations around the globe are debinding rare/unique books. I'm involved with an organization considering (not strongly) that exactly.
Internet Archive scans their books one page at a time keeping the book intact and stored in a warehouse.
usef- 53 minutes ago [-]
They do: the judge in the Anthropic case said that destroying them means that the text is being transferred, and so therefore can be fair use.
> "The print original was destroyed. One replaced the other."
(I believe the reasoning is that the author is not being financially disadvantaged by more copies of their book existing).
So the current law prefers them to destroy.
golem14 34 minutes ago [-]
Rare books seldom are under copyright and that’s not a concern then.
odyssey7 2 hours ago [-]
Once something goes commercial, academics have to consider what progress would be research-worthy, rather than a half-baked prototype for a product or feature.
Companies can’t be expected to publish their confidential and proprietary information about their feature development, and academics should consider projects that would have higher impact.
If you’re taking up a seat in a PhD program tinkering with would-be feature ideas for an existing major tech company, you should really just get hired by that tech company, where the resources are abundant and the degree is not required.
raincole 2 hours ago [-]
Do startups in other fields constantly publish their research?
xtracto 8 minutes ago [-]
That was my first thought. In my experience, particularly with trading firms (the ones doing real science based prediction trading, not TA bullshit) do not publish ANYTHING at all, and haven't for 40+ years.
Think Renaissance Technologies or similar.
TrackerFF 1 hours ago [-]
I don't expect non-foundational or non-frontier startups to do, or publish much heavy hitting research. "AI" has become such a ubiquitous description, it is probably easier to find non-AI startups.
aeternum 2 hours ago [-]
The age of the dark forest is upon us
antonvs 39 minutes ago [-]
Commercial competition has always had a dark forest character - disclosing trade secrets can lead to loss of competitive edge and in the worst case, like the dark forest, death of the company.
(Not saying that's how it should be, but that's how it has been.)
ltbarcly3 2 hours ago [-]
All startups barely publish their research, because doing that would be incredibly stupid for the most part, because they want to sell the stuff they invent not give it away for free.
gowld 1 hours ago [-]
The stuff they invent isn't what's in the papers. The progress in AI is in the engineering hurdles.
ltbarcly3 59 minutes ago [-]
What you just said is not in any relationship to reality.
deadbabe 1 hours ago [-]
There is no reason to publish anything in the AI era until you have reaped as much benefit from it as you are satisfied with. "Building in public" and "Researching in public" now means someone can swoop in with an AI and copy you instantly and become your competition overnight. Keep secrets.
Even at work, I have come to realize if I simply horde my accumulated custom AI built tools and productivity boosting tricks for myself, I can make myself more competitive as an employee.
I think this is how you get hired now, not by having a good resume, but by making claims of having special processes and personal tooling design that gets massive productivity ROI.
annoyingnoob 1 hours ago [-]
Startups exist to profit, not necessarily publish.
EGreg 2 hours ago [-]
Funny, as ONE PERSON with Claude, I've been on a tear when it comes to publishing my own research:
So I know that smart people in those companies can definitely publish. In fact, a whole team should probably be publishing like no tomorrow!
gowld 1 hours ago [-]
Have you built the things described in those papers?
EGreg 1 hours ago [-]
Yes, actually.
To be fair — from about half of the papers.
amazingamazing 2 hours ago [-]
Trade secrets are back, baby!
More to the point, without determining how much work is “worthy” of a paper it is unclear how much this matters.
Most AI companies are either a product and marketing layer over a model or not meaningfully moving any dimension to be worthy of a paper.
Also, formal papers and blogs and “cards” are all being intertwined.
sublinear 2 hours ago [-]
> “If we were racing forward on cancer-curing AI, I would be like, ’Fantastic, full steam ahead,’” she says. “But that’s not what we’re racing toward, right?”
We're not racing towards anything. We've been going in circles for years.
datakan 2 hours ago [-]
Feels more like 3 steps forward and 2 steps back.
warkdarrior 2 hours ago [-]
> We're not racing towards anything. We've been going in circles for years.
We reached the NASCAR-racing equivalent of scientific research.
xyst 2 hours ago [-]
That’s because most of its slop, and not reproducible
Revanche1367 2 hours ago [-]
Live by the non-determinism, die by the non-determinism.
kekku 2 hours ago [-]
[flagged]
semiinfinitely 2 hours ago [-]
gate keepers of antiquated practice decry people realizing they they dont need to get past the gate
The first tried to publish novel results for 3 years in tier 1 journals before finally doing a preprint and telling the tier one publishers to jump in a fire.
The second, and ongoing, isn't publishing anything because of my experience with the first.
That and avoiding openAI and Anthropic copying our results and leaving us with nothing to show for six months of work. The papers only come with the pitch deck.
In the paper, OpenAI is at the top of the chart for cumulative citations. MEGVII, Hugging Face, Waymo, Momenta, Preferred Netowkrs, Anthropic, Owkin, and Databricks, and Aibee follow (in that order). Yes, that is citations, not publications, but they explain that they're trying to use that as a proxy for significance, albeit an imperfect one.
Companies like Google aren't included because they aren't unicorn startups.
what makes them a top ai startup
We already arrived at that destination a decade ago.
Probably even further back tbh.
There's just a lot more people playing the game now, without the social indoctrination that made it more tolerable in some circles.
It's not that bad, though. September is annoying but you kind of miss the eternal renewal once you're out.
Traditional research publications are no angels. They gatekeep research, and also allow financial incentives to drive them to publish junk with their stamp on it.
But a flood of papers doesn’t actually mean more knowledge. In bypassing this route entirely, AI has swiftly lost the ability to engage with itself as a field. And the cost of that is only beginning to be felt.
The main difference is reproducibility and peer review process.
On the first, I mostly don’t care.
On the second, that’s mostly unavoidable if they want to keep their IP.
They are not perfect, but institutions like peer review or universities do provide a framework and incentives for quality.
https://research.google/pubs/distilling-the-knowledge-in-a-n...
# LLM generated summary of the implied irony
Google may consider the standalone frontier-model arms race economically irrational, while still considering frontier-model capability strategically indispensable. Its longer game is probably not to avoid building the biggest models, but to build only enough of them to serve as capability factories—then turn that intelligence into a much larger population of cheap, purpose-built models.
(Human again) If Google knew what they were on to, why wouldn’t they make it their secret weapon from the start? I suspect it’s because they predicted there would be an arms race, and knew how to profit from it. They had a distillation paper published before “Attention is all you need”. In hindsight, is it ironic at all? Or is it obvious?
It was kind of like radiation science before WWII, it would be freely published because it wasn't potentially world changing yet. After it became a government interest, even people doing things unrelated to weapons would be much more apt to hold their work close.
50% of startups contributing to public research is actually a crazy good outcome. That’s far more than I had expected.
There should be a new ESG (Environmental, Social, and Governance) policy being pushed recognizing the important role this plays. Although ESG and all norms have been set aside in this grim new world it seems.
You are not. A disproportionate amount of value in the computing industry was created by <strike>geniuses</strike> decently smart people who worked together and who decided to just tell people how to do things instead of trying to capture the value of being the first person to figure out how to do those things.
This observation pre-dates the current wave of AI hype by a half century or so.
> driven by greed
I can only speak for myself.
For me it's exactly the opposite. If I want to explain how something works, I can just... do that. If I want to share an artifact demonstrating how to solve a particular type of problem, I can just... do that. If I want to mentor/teach, I can just... do that.
Doing those things within the confines of Academia Approved Institutions is exhausting and distracting.
To wit, and the actual point of this post: the term "Publishing Research" in this article doesn't mean "post it on a .html page and share the source code". It means engaging in a very specific and peculiar and extremely political modality of communication.
And it really only makes sense to do that specific and peculiar and political thing you're at a stage in your professional/personal development where you need to play that particular prestige game. (Which there's nothing wrong with, but it is a deeply cargo culted version of the actual scientific process.)
* A disproportionate amount of value in the computing industry was created by g̶e̶n̶i̶u̶s̶e̶s̶ decently smart people who worked together to do things that seemed impossible and who decided to just tell people how to do things instead of trying to capture the value of being the first person to figure out how to do those things.*
edit: edited.
Fortunately
There is some excellent publicly-funded research in there, but pivotal papers like Attention Is All You Need and the numerous pivotal OpenAI publications were privately funded.
OpenAI is at the top of the chart in the study.
I think you're bringing some assumptions into this conversation that aren't supported by the evidence.
A distinction should be made between the old nonprofit OpenAI and the current organization. They don't publish technical research anymore.
Like I said, I think you're bringing some assumptions to this conversation that aren't based on the how the industry came about.
Example: https://arxiv.org/abs/2607.24653
Exceptions to some like Deepmind, Nvidia, and Thinking Machines
[1] https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
The discussion has been around, but it's flared significantly recently. Not sure if that's what the person you're responding to is specifically inflamed about.
Regardless, old/rare books are certainly being acquired and destroyed.
Elon tweeted "I’ve asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning"
which is no sense a credible source for what he actually asked SpaceXAI team to do, or if they are doing it, or what they were doing before yesterday. Elon has an extremely long and thorough track record of lying.
https://xcancel.com/elonmusk/status/2081844165881594362
The problem remains either way. Small operations around the globe are debinding rare/unique books. I'm involved with an organization considering (not strongly) that exactly.
The scanning process destroys the book.
Internet Archive scans their books one page at a time keeping the book intact and stored in a warehouse.
> "The print original was destroyed. One replaced the other."
(I believe the reasoning is that the author is not being financially disadvantaged by more copies of their book existing).
So the current law prefers them to destroy.
Companies can’t be expected to publish their confidential and proprietary information about their feature development, and academics should consider projects that would have higher impact.
If you’re taking up a seat in a PhD program tinkering with would-be feature ideas for an existing major tech company, you should really just get hired by that tech company, where the resources are abundant and the degree is not required.
Think Renaissance Technologies or similar.
(Not saying that's how it should be, but that's how it has been.)
Even at work, I have come to realize if I simply horde my accumulated custom AI built tools and productivity boosting tricks for myself, I can make myself more competitive as an employee.
I think this is how you get hired now, not by having a good resume, but by making claims of having special processes and personal tooling design that gets massive productivity ROI.
https://arxiv.org/search/?searchtype=author&query=Magarshak%...
So I know that smart people in those companies can definitely publish. In fact, a whole team should probably be publishing like no tomorrow!
To be fair — from about half of the papers.
More to the point, without determining how much work is “worthy” of a paper it is unclear how much this matters.
Most AI companies are either a product and marketing layer over a model or not meaningfully moving any dimension to be worthy of a paper.
Also, formal papers and blogs and “cards” are all being intertwined.
We're not racing towards anything. We've been going in circles for years.
We reached the NASCAR-racing equivalent of scientific research.