Category: Science
My excellent Conversation with Sam Altman
Recorded live in Berkeley, at the Roots of Progress conference (an amazing event), here is the material with transcript, here is the episode summary:
Sam Altman makes his second appearance on the show to discuss how he’s managing OpenAI’s explosive growth, what he’s learned about hiring hardware people, what makes roon special, how far they are from an AI-driven replacement to Slack, what GPT-6 might enable for scientific research, when we’ll see entire divisions of companies run mostly by AI, what he looks for in hires to gauge their AI-resistance, how OpenAI is thinking about commerce, whether GPT-6 will write great poetry, why energy is the binding constraint to chip-building and where it’ll come from, his updated plan for how he’d revitalize St. Louis, why he’s not worried about teaching normies to use AI, what will happen to the price of healthcare and hosing, his evolving views on freedom of expression, why accidental AI persuasion worries him more than intentional takeover, the question he posed to the Dalai Lama about superintelligence, and more.
Excerpt:
COWEN: What is it about GPT-6 that makes that special to you?
ALTMAN: If GPT-3 was the first moment where you saw a glimmer of something that felt like the spiritual Turing test getting passed, GPT-5 is the first moment where you see a glimmer of AI doing new science. It’s very tiny things, but here and there someone’s posting like, “Oh, it figured this thing out,” or “Oh, it came up with this new idea,” or “Oh, it was a useful collaborator on this paper.” There is a chance that GPT-6 will be a GPT-3 to 4-like leap that happened for Turing test-like stuff for science, where 5 has these tiny glimmers and 6 can really do it.
COWEN: Let’s say I run a science lab, and I know GPT-6 is coming. What should I be doing now to prepare for that?
ALTMAN: It’s always a very hard question. Even if you know this thing is coming, if you adapt your —
COWEN: Let’s say I even had it now, right? What exactly would I do the next morning?
ALTMAN: I guess the first thing you would do is just type in the current research questions you’re struggling with, and maybe it’ll say, “Here’s an idea,” or “Run this experiment,” or “Go do this other thing.”
COWEN: If I’m thinking about restructuring an entire organization to have GPT-6 or 7 or whatever at the center of it, what is it I should be doing organizationally, rather than just having all my top people use it as add-ons to their current stock of knowledge?
ALTMAN: I’ve thought about this more for the context of companies than scientists, just because I understand that better. I think it’s a very important question. Right now, I have met some orgs that are really saying, “Okay, we’re going to adopt AI and let AI do this.” I’m very interested in this, because shame on me if OpenAI is not the first big company run by an AI CEO, right?
COWEN: Just parts of it. Not the whole thing.
ALTMAN: No, the whole thing.
COWEN: That’s very ambitious. Just the finance department, whatever.
ALTMAN: Well, but eventually it should get to the whole thing, right? So we can use this and then try to work backwards from that. I find this a very interesting thought experiment of what would have to happen for an AI CEO to be able to do a much better job of running OpenAI than me, which clearly will happen someday. How can we accelerate that? What’s in the way of that? I have found that to be a super useful thought experiment for how we design our org over time and what the other pieces and roadblocks will be. I assume someone running a science lab should try to think the same way, and they’ll come to different conclusions.
COWEN: How far off do you think it is that just, say, one division of OpenAI is 85 percent run by AIs?
ALTMAN: Any single division?
COWEN: Not a tiny, insignificant division, mostly run by the AIs.
ALTMAN: Some small single-digit number of years, not very far. When do you think I can be like, “Okay, Mr. AI CEO, you take over”?
Of course we discuss roon as well, not to mention life on the moons of Saturn…
Track the 2025-2026 Economics Job Market
“Econ Now aggregates economics PhD job market papers, job postings for economists, and economics conference information all in one place.”
Link here, I have yet to try it.
Supply is elastic, in a new and different setting
The longstanding debate over whether human capabilities and skills are shaped more by “nature” or “nurture” has been revitalized by recent advances in genetics, particularly in the use of polygenic scores (PGSs) to proxy for genetic endowments. Yet, we argue that PGSs embed not only direct genetic effects but also indirect environmental influences, raising questions about their validity for causal analysis. We show that these conflated measures can mislead studies of gene–environment interactions, especially when parental behavior responds to children’s genetic risk. To address this issue, we construct a new latent measure of genetic risk that integrates individual genotypes with diagnostic symptoms, using data from the National Longitudinal Study of Adolescent to Adult Health linked to restricted individual SNP-level genotypes from dbGaP. Exploiting multiple sources of variation—including the Mendelian within-family genetic randomization among siblings—we find consistent evidence that parents compensate by investing more in children with higher genetic risk for ADHD. Strikingly, these compensatory responses disappear when genetic risk is proxied by the conventional ADHD PGS, which also yields weaker—and in some cases reversed—predictions for long-run outcomes. Finally, we embed our latent measure of genetic endowments into a standard dynamic structural model of child development. The model shows that both parental investments and latent genetic risk jointly shape children’s cognitive and mental health development, underscoring the importance of modeling the dynamic interplay between genes and environments in the formation of human capital.
That is from a new NBER working paper, by
Observations on browsing economics job market candidates
The number of people on the market seems much lower this year, perhaps because of the lag with Covid, as well as more general demographic trends. Even adjusting for the lower number of candidates, I found fewer interesting papers this year than usual, as research interests continue to narrow. There is too much emphasis on showing quality technique by answering a small question well, rather than addressing more important questions more imperfectly. Harvard had by far the most interesting students, as most of them were considering questions I cared about. LSE looked pretty good too. In terms of topics, I saw a lot of papers on educational testing, urban economics and mobility, and AI. Theory seems to be permanently on the wane. The number of co-authors continues to rise.
Overall I came away with a bad feeling from this year’s perusal, noting there are some departments I have not looked at yet. In the aggregate it did not seem vital enough or exciting enough to me?
I still will be putting up some more of the papers I found of interest.
A new critique of RCTs
Randomized Controlled Trials (RCTs) are the gold standard for evaluating the effects of interventions because they rely on simple assumptions. Their validity also depends on an implicit assumption: that the research process itself—including how participants are assigned—does not affect outcomes. In this paper, I challenge this assumption by showing that outcomes can depend on the subject’s knowledge of the study, their treatment status, and the assignment mechanism. I design a field experiment in India around a soil testing program that exogenously varies how participants are informed of their assignment. Villages are randomized into two main arms: one where treatment status is determined by a public lottery, and another by a private computerized process. My design temporally separates assignment from treatment delivery, allowing me to isolate the causal effect of the assignment process itself. I find that estimated treatment effects differ across assignment methods and that these effects emerge even before the treatment is delivered. The effects are not uniform: the control group responds more strongly to the assignment method than the treated group. These findings suggest that the choice of assignment procedure is consequential and that failing to account for it can threaten the interpretation and generalizability of standard RCT treatment effect estimates.
That is the job market paper of Florencia Hnilo, from Stanford, who also does economic history,
Some further new negative results on minimum wage hikes
We study how exposure to scientific research in university laboratories influences students’ pursuit of careers in science. Using administrative data from thousands of research labs linked to student career outcomes and a difference-in-differences design, we show that state minimum wage increases reduce employment of undergraduate research assistants in labs by 7.4%. Undergraduates exposed to these minimum wage increases graduate with 18.1% fewer quarters of lab experience. Using minimum wage changes as an instrumental variable, we estimate that one fewer quarter working in a lab, particularly early in college, reduces the probability of working in the life sciences industry by 2 percentage points and of pursuing doctoral education by 7 percentage points. These effects are attenuated for students supported by the Federal Work-Study program. Our findings highlight how labor market policies can shape the career paths of future scientists and the importance of budget flexibility for principal investigators providing undergraduates with research experience.
That is from a new working paper by Ina Ganguli and Raviv Murciano-Goroff.
When will quantum computing work?
Huge investments are flowing into QC companies today. IonQ has a $19B market cap, Rigetti has a $10B cap, and PsiQuantum recently raised $1B.3D-Wave is not relevant, despite high qubit counts. Their machines are annealers, rather than gate based, and have less computational power than the QCs that IonQ, Rigetti, PsiQuantum, etc. are working on. This is a lot of money for an industry generating no real revenue, and without an apparent path to revenue over the next 5 years. Qubit counts have not been doubling each year, but even if they did, we’d have 32 kq machines in 2030.4If qubits double each year, 1,000 qubits today grows to 32 kq in 5 years’ time. There are few – if any – commercial applications for machines of that size. Will these companies keep raising larger rounds until they achieve 100 kq? Or have they got some secret sauce we don’t know about that investors are betting on? If there has been a true breakthrough, we should see much faster growth in qubit count, as well as larger and larger quantum processors, running increasingly massive programs. Note that the QC ecosystem is reasonably public and both private companies and university labs are competitive players. Advances tend to get published rather than stowed away.
Here is more from Tom McCarthy.
Understanding and Addressing Temperature Impacts on Mortality
Here are some important results:
A large literature documents how ambient temperature affects human mortality. Using decades of detailed data from 30 countries, we revisit and synthesize key findings from this literature. We confirm that ambient temperature is among the largest external threats to human health, and is responsible for a remarkable 5-12% of total deaths across countries in our sample, or hundreds of thousands of deaths per year in both the U.S. and EU. In all contexts we consider, cold kills more than heat, though the temperature of minimum risk rises with age, making younger individuals more vulnerable to heat and older individuals more vulnerable to cold. We find evidence for adaptation to the local climate, with hotter places experiencing somewhat lower risk at higher temperatures, but still more overall mortality from heat due to more frequent exposure. Within countries, higher income is not associated with uniformly lower vulnerability to ambient temperature, and the overall burden of mortality from ambient temperature is not falling over time. Finally, we systematically summarize the limited set of studies that rigorously evaluate interventions that can reduce the impact of heat and cold on health. We find that many proposed and implemented policy interventions lack empirical support and do not target temperature exposures that generate the highest health burden, and that some of the most beneficial interventions for reducing the health impacts of cold or heat have little explicit to do with climate.
Those are from a recent paper by Marshall Burke, et.al.
We Turned the Light On—and the AI Looked Back
Jack Clark, Co-founder of Anthropic, has written a remarkable essay about his fears and hopes. It’s not the usual kind of thing one reads from a tech leader:
I remember being a child and after the lights turned out I would look around my bedroom and I would see shapes in the darkness and I would become afraid – afraid these shapes were creatures I did not understand that wanted to do me harm. And so I’d turn my light on. And when I turned the light on I would be relieved because the creatures turned out to be a pile of clothes on a chair, or a bookshelf, or a lampshade.
Now, in the year of 2025, we are the child from that story and the room is our planet. But when we turn the light on we find ourselves gazing upon true creatures, in the form of the powerful and somewhat unpredictable AI systems of today and those that are to come. And there are many people who desperately want to believe that these creatures are nothing but a pile of clothes on a chair, or a bookshelf, or a lampshade. And they want to get us to turn the light off and go back to sleep.
…We are growing extremely powerful systems that we do not fully understand. Each time we grow a larger system, we run tests on it. The tests show the system is much more capable at things which are economically useful. And the bigger and more complicated you make these systems, the more they seem to display awareness that they are things.
It is as if you are making hammers in a hammer factory and one day the hammer that comes off the line says, “I am a hammer, how interesting!” This is very unusual!
…I am also deeply afraid. It would be extraordinarily arrogant to think working with a technology like this would be easy or simple.
My own experience is that as these AI systems get smarter and smarter, they develop more and more complicated goals. When these goals aren’t absolutely aligned with both our preferences and the right context, the AI systems will behave strangely.
…we are not yet at “self-improving AI”, but we are at the stage of “AI that improves bits of the next AI, with increasing autonomy and agency”. And a couple of years ago we were at “AI that marginally speeds up coders”, and a couple of years before that we were at “AI is useless for AI development”. Where will we be one or two years from now?
And let me remind us all that the system which is now beginning to design its successor is also increasingly self-aware and therefore will surely eventually be prone to thinking, independently of us, about how it might want to be designed.
…In closing, I should state clearly that I love the world and I love humanity. I feel a lot of responsibility for the role of myself and my company here. And though I am a little frightened, I experience joy and optimism at the attention of so many people to this problem, and the earnestness with which I believe we will work together to get to a solution. I believe we have turned the light on and we can demand it be kept on, and that we have the courage to see things as they are.
Clark is clear that we are growing intelligent systems that are more complex than we can understand. Moreover, these systems are becoming self-aware–that is a fact, even if you think they are not sentient (but beware hubris on the latter question).
Nobel Prize in economics goes to Philippe Aghion and Peter Howitt and Joel Mokyr
Excellent choices. Here is the press release with links to longer discussions of their works.
This is a prize for economic growth, and for the ideas of creative destruction. Those are some of the most important ideas in economics, so I could not be happier with this pick. Joel Mokyr in particular has been a long-time associate of GMU and Mercatus, so I would like to congratulate him in particular.
Aghion is at INSEAD in France, Howitt at Brown, and Mokyr at Northwestern. It is also nice to see some people outside of “the usual schools” winning the prize. Aghion and Howitt, of course, worked together to produce a model of creative destruction and economic growth. Here are their key papers together.
Joel Mokyr is an economic historian, and best known for his pioneering work in explaining the Industrial Revolution in England. Here are his best-known works. Read The Lever of Riches and The Gifts of Athena and A Culture of Growth. I have benefited most from The Enlightened Economy: An Economic History of Britain 1700-1850. He has a new book coming out in November, with Tabellini and Greif. It is correct to consider him as an “Enlightenment thinker.” Brian Albrecht has a good thread on this.
Below you can find individual posts on Aghion, Howitt, and also Mokyr. Here is Alex’s post on the prize.
Science Policy Insider
That is a new Substack by Jim Olds, here is the introduction:
How science funding really works—from someone who ran the machinery at NSF and NIH.
I’m Jim Olds, former head of the National Science Foundation’s $750M Biological Sciences Directorate (2014-2018), NSF lead for President Obama’s BRAIN Initiative, and co-chair of the White House Life Sciences Subcommittee.
Over three decades in Washington D.C., I’ve seen how science policy actually works—not from the sidelines, but from inside the decision-making rooms. I’ve reviewed thousands of grants, managed billion-dollar budgets, and worked with everyone from Nobel laureates to members of Congress.
What you’ll get here:
– The real story of how funding decisions get made
– Insider analysis of science policy debates and initiatives
– Practical insights on what makes big science succeed or fail
– Honest perspectives on the challenges facing research funding
This isn’t speculation or critique from the outside. It’s the view from someone who was in the room where it happened.
This newsletter is free
What should I ask Blake Scholl?
The Boom guy doing supersonic flight! I will be doing a podcast with him at the forthcoming Roots of Progress Institute event. He already has told “his story” on the supersonic side a number of times, so what else of interest should we cover? For background, here is Blake on Wikipedia.
New archaeology tranche for Emergent Ventures
Just apply at the normal site. Here is a description of what we are up to and what we are looking for:
- We are giving archaeology-specific EV grants
- Emphasis on projects enabled by tech (AI/CV, lidar, synthetic aperture radar and hyperspectral imagery, open source data, etc)
- Flexible on cost or duration of projects
- All circumstances encouraged (grad student funding a project she’s working on, swe wanting to take a few months off to work on something, high school student side project, etc)
- Ideas below are by no means comprehensive or the edges of the search area – just starting points for exploration for people interested in the field that don’t have a specific idea in mind yet
Examples of ideas:
- Transcribing, digitizing, and open-sourcing cuneiform tablets and other writing
- OCR and cataloguing old scripts (Maya glyphs, Sanskrit, etc)
- Scanning coastlines along old civilizations for submerged sites
- Improved ground penetrating radar
- Improving AI for translation of old and obscure languages, and mass translation
- Deciphering Linear A. Are existing proposed decipherments such as this one plausible? https://x.com/natfriedman/status/1756414276779253872?s=46
- Reading lumped Maya codices (https://en.wikipedia.org/wiki/Maya_codices#Other_Maya_codices)
- Reading Chinese bamboo slips and silk manuscripts (https://x.com/jordanschnyc/status/1920181694327362007)
- Using AI to check vast amounts of land for sites (good example: https://x.com/byornoste/status/1927434033685827726)
- Lidar scanning forests to find structures beneath the canopy (https://x.com/mehran__jalali/status/1934693744630223228)
- Using SAR/hyperspectral scans for finding sites
For these notes I thank Mehran Jalali, a former EV winner in this area, who also will serve as one of the referees. When you apply, just indicate that your request is for archaeology. Soon this will be a formal category on the application itself, if somehow you are already ready to apply tomorrow a.m., just use the word in your project description.
We thank Yonatan Ben Shimon for his generous support of this tranche.
AI Scientists in the Lab
Today, we introduce Periodic Labs. Our goal is to create an AI scientist.
Science works by conjecturing how the world might be, running experiments, and learning from the results.
Intelligence is necessary, but not sufficient. New knowledge is created when ideas are found to be consistent with reality. And so, at Periodic, we are building AI scientists and the autonomous laboratories for them to operate.
…Autonomous labs are central to our strategy. They provide huge amounts of high-quality data (each experiment can produce GBs of data!) that exists nowhere else. They generate valuable negative results which are seldom published. But most importantly, they give our AI scientists the tools to act.
…One of our goals is to discover superconductors that work at higher temperatures than today’s materials. Significant advances could help us create next-generation transportation and build power grids with minimal losses. But this is just one example — if we can automate materials design, we have the potential to accelerate Moore’s Law, space travel, and nuclear fusion.
Our founding team co-created ChatGPT, DeepMind’s GNoME, OpenAI’s Operator (now Agent), the neural attention mechanism, MatterGen; have scaled autonomous physics labs; and have contributed to important materials discoveries of the last decade. We’ve come together to scale up and reimagine how science is done.
The AI’s can work 24 hours a day, 365 days a year and with labs under their control the feedback will be quick. In nine hours, AlphaZero taught itself chess and then trounced the then world champion Stockfish 8, (ELO around 3378 compared to Magnus Carlsen’s high of 2882). That was in 2017. In general, experiments are more open-ended than chess but not necessarily in every domain. Moreover context windows and capabilities have grown tremendously since 2017.
In other AI news, AI can be used to generate dangerous proteins like ricin and current safeguards are not very effective:
Microsoft bioengineer Bruce Wittmann normally uses artificial intelligence (AI) to design proteins that could help fight disease or grow food. But last year, he used AI tools like a would-be bioterrorist: creating digital blueprints for proteins that could mimic deadly poisons and toxins such as ricin, botulinum, and Shiga.
Wittmann and his Microsoft colleagues wanted to know what would happen if they ordered the DNA sequences that code for these proteins from companies that synthesize nucleic acids. Borrowing a military term, the researchers called it a “red team” exercise, looking for weaknesses in biosecurity practices in the protein engineering pipeline.
The effort grew into a collaboration with many biosecurity experts, and according to their new paper, published today in Science, one key guardrail failed. DNA vendors typically use screening software to flag sequences that might be used to cause harm. But the researchers report that this software failed to catch many of their AI-designed genes—one tool missed more than 75% of the potential toxins.
Solve for the equilibrium?
Good job people, congratulations…
“Sonnet 4.5 does complete replication checks of an econpaper.”
That is Kevin Bryan, here is more from Ethan Mollick.