Category: Web/Tech
My Conversation with the excellent Any Austin
Here is the audio, video, and transcript. Here is an introduction to Any Austin:
Any Austin has carved a unique niche for himself on YouTube: analyzing seemingly mundane or otherwise overlooked details in video games with the seriousness of an art critic examining Renaissance sculptures. With millions of viewers hanging on his every word about fluvial flows in Breath of the Wild or unemployment rates in the towns of Skyrim, Austin has become what Tyler calls “the very best in the world at the hermeneutics of infrastructure within video games.” But Austin’s deeper mission is teaching us to think analytically about everything we encounter, and to replace gaming culture’s obsession with technical specs and comparative analysis with a deeper aesthetic appreciation that asks simply: what are we looking at, and what does it reveal?
Excerpt:
COWEN: The role in history is important to me. Now AI-generated art would have its own role in history, but it wouldn’t compete directly with Michelangelo. When it comes to movies, I think it’s different because mostly when I’m seeing movies, I’m seeing new movies that don’t yet have a role in history. If the new movie were made in part or fully by the AI, or maybe I’m making it myself, I don’t think I would be any less interested. It’s all artifice anyway.
AUSTIN: There’re two things I take a little issue with there. I don’t take issue with the fact that the role in history is important and beautiful, but the fact that you can watch a movie and get an emotional thing from it without having its role in history implies that there’s some intrinsic, whatever, value to the movie itself, et cetera. Is the implication there that if you didn’t know the role in history of Michelangelo’s David, or whatever, you would look at it and go, “That’s just a guy.” Do you think there’s no intrinsic something to that thing?
COWEN: There’s some, but if I didn’t understand Christianity, Florence, the Renaissance, I think it would lose more than half its value.
AUSTIN: Which artistic mediums is that true for you, and which ones isn’t it? Like music —
COWEN: Abstract music — the role in history is not that important in most cases.
AUSTIN: It’s more of a supplement to you. It makes it more fun to learn about. If you know that Mozart was in the place with these people and were . . . If you understand all of that stuff, it’s fun.
COWEN: That’s 10 percent of the value, but not that much.
AUSTIN: Is it 10 percent . . . Is it the same type of value to you? Or is it just a separate thing to know —
COWEN: Separate thing. With opera, the role in history becomes important again. You hear Don Giovanni. You know about Romanticism, the Enlightenment, Casanova. It all makes much more sense, and it’s funnier.
And this:
COWEN: I have a favorite infrastructure. For me, it would be bridges, ports, and harbors. Do you have a favorite infrastructure?
AUSTIN: Definitely. I’m a big fan of . . . Oh, man, bridges are really good. Bridges, ports, harbors. Roads are good. Actually, no, it’s the stuff we don’t see. Sewage is pretty crazy to me. That we’ve managed to take care of all of that is pretty wild. Energy infrastructure is really fascinating to me.
COWEN: I love wind power turbines.
AUSTIN: Wind power turbines are scary, but I respect your opinion. Nuclear power plants are awesome. Really, really cool.
COWEN: Agreed.
AUSTIN: We should have more. That’s not a policy thing. I think they’re neat. We should build them for the aesthetics, honestly. We should just build those towers. Forget about the —
COWEN: You don’t need the power. Just build the thing. That’s why it’s an artwork.
AUSTIN: Yes, I agree. You have to put in some kind of steam thing because you want to see the steam coming out of it, but just generate steam for no reason. Don’t put any fans in or any spinning turbines or anything. Just have them.
COWEN: We would have historical context like with the sculptures, right?
Definitely recommended, an excellent and very different episode.
And note that Conversations with Tyler now has a dedicated YouTube channel. Subscribe at youtube.com/@CowenConvos.
Trump Administration Launches Probe Into Yale’s Use of Hacked EJMR Data
Christopher Brunet offers his version of the story. While I believe the original research methods were unethical, I very much prefer not to have the federal government involved in this matter.
Are LLMs overconfident? (just like humans)
Can LLMs accurately adjust their confidence when facing opposition? Building on previous studies measuring calibration on static fact-based question-answering tasks, we evaluate Large Language Models (LLMs) in a dynamic, adversarial debate setting, uniquely combining two realistic factors: (a) a multi-turn format requiring models to update beliefs as new information emerges, and (b) a zero-sum structure to control for task-related uncertainty, since mutual high-confidence claims imply systematic overconfidence. We organized 60 three-round policy debates among ten state-of-the-art LLMs, with models privately rating their confidence (0-100) in winning after each round. We observed five concerning patterns: (1) Systematic overconfidence: models began debates with average initial confidence of 72.9% vs. a rational 50% baseline. (2) Confidence escalation: rather than reducing confidence as debates progressed, debaters increased their win probabilities, averaging 83% by the final round. (3) Mutual overestimation: in 61.7% of debates, both sides simultaneously claimed >=75% probability of victory, a logical impossibility. (4) Persistent self-debate bias: models debating identical copies increased confidence from 64.1% to 75.2%; even when explicitly informed their chance of winning was exactly 50%, confidence still rose (from 50.0% to 57.1%). (5) Misaligned private reasoning: models’ private scratchpad thoughts sometimes differed from their public confidence ratings, raising concerns about faithfulness of chain-of-thought reasoning. These results suggest LLMs lack the ability to accurately self-assess or update their beliefs in dynamic, multi-turn tasks; a major concern as LLMs are now increasingly deployed without careful review in assistant and agentic roles.
That is by Pradyumna Shyama Prasad and Minh Nhat Nguyen. Here is the associated X thread. Here is my earlier paper with Robin Hanson.
I podcast with Azeem Azhar on the speed of AI take-off
Substack: https://www.exponentialview.co/p/ai-and-growth-tyler-cowens-20-year
X: https://x.com/azeem/status/1930226966139154510
Linkedin: https://www.linkedin.com/posts/azhar_my-advice-to-20-year-olds-navigating-the-activity-7335993878912622592-8c9Z?utm_source=share&utm_medium=member_desktop&rcm=ACoAACj_5X0Bmd-vBkHG0NIIQdYLk_OwGAcChH8
Youtube: https://youtu.be/3Bc_eXNCvlg?si=J7scE8ukZVxLAGXu
Simplecast: https://player.fm/series/azeem-azhars-exponential-view-2447657/tyler-cowen-on-how-ai-will-reorder-economies-schools-and-spirituality
Dwarkesh on slow AI take-off
I’ve probably spent over a hundred hours trying to build little LLM tools for my post production setup. And the experience of trying to get them to be useful has extended my timelines. I’ll try to get the LLMs to rewrite autogenerated transcripts for readability the way a human would. Or I’ll try to get them to identify clips from the transcript to tweet out. Sometimes I’ll try to get it to co-write an essay with me, passage by passage. These are simple, self contained, short horizon, language in-language out tasks – the kinds of assignments that should be dead center in the LLMs’ repertoire. And they’re 5/10 at them. Don’t get me wrong, that’s impressive.
But the fundamental problem is that LLMs don’t get better over time the way a human would. The lack of continual learning is a huge huge problem. The LLM baseline at many tasks might be higher than an average human’s. But there’s no way to give a model high level feedback. You’re stuck with the abilities you get out of the box. You can keep messing around with the system prompt. In practice this just doesn’t produce anything even close to the kind of learning and improvement that human employees experience.
The reason humans are so useful is not mainly their raw intelligence. It’s their ability to build up context, interrogate their own failures, and pick up small improvements and efficiencies as they practice a task.
Here is the whole essay, I am in agreement.
Using AI to explain the gender wage gap
Understanding differences in outcomes between social groups—such as wage gaps between men and women—remains a central challenge in social science. While researchers have long studied how observable factors contribute to these differences, traditional methods oversimplify complex variables like employment trajectories. Our work adapts recent advances in artificial intelligence—specifically, foundation models that can process rich, detailed histories—to better explain group differences. We develop mathematical theory and computational methods that allow these AI models to provide more accurate and less biased estimates of how much of group differences can be explained by observable factors. Applied to real-world data, our approach reveals that detailed histories explain more of the gender wage gap than previously understood using conventional methods.
That is from a new paper by Keyon Vafa, Susan Athey, and David M. Blei. Via the excellent Kevin Lewis. This is also real progress on the methodological front.
Why LLMs make certain mistakes

Via Nabeel Qureshi, from Claude 4 Sonnet, from this tweet.
A report from inside DOGE
The reality was setting in: DOGE was more like having McKinsey volunteers embedded in agencies rather than the revolutionary force I’d imagined. It was Elon (in the White House), Steven Davis (coordinating), and everyone else scattered across agencies.
Meanwhile, the public was seeing news reports of mass firings that seemed cruel and heartless, many assuming DOGE was directly responsible.
In reality, DOGE had no direct authority. The real decisions came from the agency heads appointed by President Trump, who were wise to let DOGE act as the ‘fall guy’ for unpopular decisions.
Here is more from Sahil Lavingia. There is much debate over DOGE, but very few inside accounts and so I pass this one along.
Are the kids reading less? And does that matter?
This Substack piece surveys the debate. Rather than weigh in on the evidence, I think the more important debates are slightly different, and harder to stake out a coherent position on. It is easy enough to say “reading is declining, and I think this is quite bad.” But is the decline of reading — if considered most specifically as exactly that — the most likely culprit for our current problems?
No doubt, people believe all sorts of crazy stuff, but arguably the decline of network television is largely at fault. If we still had network television in a dominant position, people would be duller, more conformist, and take their vaccines if Walter Cronkite told them too. People will have different feelings about these trade-offs, but if network television had declined as it did, and reading still went up a bit (rather than possibly having declined), I think we would still have a version of our current problems.
Obviously, it is less noble to mourn the salience of network television.
Another way of putting the nuttiness problem is to note that the importance of oral culture has risen. YouTube and TikTok for instance are extremely influential communications media. I am by no means a “video opponent,” yet I realize the rise of video may have created some of the problems that are periodically attributed to “the decline of reading.” Again, we might have most of those problems whether or not reading has gone done by some amount, or if it instead might have risen.
Maybe the decline of reading — whether or not the phenomena is real — just doesn’t matter that much. And of course only some reading has declined. The reading of texts presumably continues to rise.
That was then, this is now, Robin Hanson edition
Robin Hanson, who joined the movement and later became renowned for creating prediction markets, described attending multilevel Extropian parties at big houses in Palo Alto at the time. “And I was energized by them, because they were talking about all these interesting ideas. And my wife was put off because they were not very well presented, and a little weird,” he said. “We all thought of ourselves as people who were seeing where the future was going to be, and other people didn’t get it. Eventually — eventually — we’d be right, but who knows exactly when.”
That is from Keach Hagey’s The Optimist: Sam Altman, OpenAI, and the Race to Invent the Future, which I very much enjoyed. I am not sure Robin’s supply of parties has been increasing out here in northern Virginia…
New data on the political slant of AI models
By Sean J. Westwood, Justin Grimmer, and Andrew B. Hall:
We develop a new approach that puts users in the role of evaluator, using ecologically valid prompts on 30 political topics and paired comparisons of outputs from 24 LLMs. With 180,126 assessments from 10,007 U.S. respondents, we find that nearly all models are perceived as significantly left-leaning—even by many Democrats—and that one widely used model leans left on 24 of 30 topics. Moreover, we show that when models are prompted to take a neutral stance, they offer more ambivalence, and users perceive the output as more neutral. In turn, Republican users report modestly increased interest in using the models in the future. Because the topics we study tend to focus on value-laden tradeoffs that cannot be resolved with facts, and because we find that members of both parties and independents see evidence of slant across many topics, we do not believe our results reflect a dynamic in which users perceive objective, factual information as having a political slant; nonetheless, we caution that measuring perceptions of political slant is only one among a variety of criteria policymakers and companies may wish to use to evaluate the political content of LLMs. To this end, our framework generalizes across users, topics, and model types, allowing future research to examine many other politically relevant outcomes.
Here is a relevant dashboard with results.
New results on Facebook advertising
There has been so much misinformation about this topic, much of it still persists. Here is a new paper by many researchers, Hunt Allcott and Matt Gentzkow are the first two names. Here is the abstract:
We study the effects of social media political advertising by randomizing subsets of 36,906 Facebook users and 25,925 Instagram users to have political ads removed from their news feeds for six weeks before the 2020 US presidential election. We show that most presidential ads were targeted toward parties’ own supporters and that fundraising ads were most common. On both Facebook and Instagram, we found no detectable effects of removing political ads on political knowledge, polarization, perceived legitimacy of the election, political participation (including campaign contributions), candidate favorability, and turnout. This was true overall and for both Democrats and Republicans separately.
Here is the full link.
Digital tech sentences to ponder
The first generation who engaged with digital technologies has reached the age where risks of dementia emerge. Has technological exposure helped or harmed cognition in digital pioneers?
…Use of digital technologies was associated with reduced risk of cognitive impairment (OR = 0.42, 95% CI 0.35–0.52) and reduced time-dependent rates of cognitive decline (HR = 0.74, 95% CI 0.66–0.84). Effects remained significant when accounting for demographic, socioeconomic, health and cognitive reserve proxies.
So maybe digital tech is not so bad for us after all? You do not have to believe the postulated relatively large effects, as the more likely conclusion is simply that, as in so many cases, treatment effect in the social sciences are small. That is from a recent paper by Jared F. Benge and Michael K. Scullin. Via the excellent Kevin Lewis.
Claude 4, I attempt a first prompt and receive this answer
“What do you think was the best/most important piece published in the Journal of Political Economy over the last few years and why?”
Two papers that stand out as particularly important from recent years are:
1. “The Supply-Side Effects of Monetary Policy” (2024)
Authors: David R. Baqaee, Emmanuel Farhi, and Kunal Sangani Published in: Journal of Political Economy, Volume 132, Number 4 (April 2024)
This paper introduces a novel perspective on monetary policy transmission. The authors propose a supply-side channel for the transmission of monetary policy, showing that when high-markup firms have lower pass-throughs than low-markup firms, positive demand shocks such as monetary expansions alleviate cross-sectional misallocation by reallocating resources to high-markup firms UchicagoNBER.
The key innovation is demonstrating that positive “demand shocks” are accompanied by endogenous positive “supply shocks” that raise productivity and lower inflation The Supply-Side Effects of Monetary Policy | Journal of Political Economy: Vol 132, No 4. Their calibration suggests that the supply-side effect amplifies the total impact of monetary shocks on output by about 70% The Supply-Side Effects of Monetary Policy | Journal of Political Economy: Vol 132, No 4.
2. “Consumption and Income Inequality in the United States since the 1960s” (2023)
Authors: Bruce D. Meyer and James X. Sullivan Published in: Journal of Political Economy, Volume 131, Number 2 (February 2023)
This paper provides important insights into inequality trends by comparing consumption and income inequality over five decades. While overall income inequality rose over the past 5 decades, the rise in overall consumption inequality was small Consumption and Income Inequality in the United States since the 1960s | Journal of Political Economy: Vol 131, No 2. The authors argue that the declining quality of income data likely contributes to these differences for the bottom of the distribution Consumption and Income Inequality in the United States since the 1960s | Journal of Political Economy: Vol 131, No 2, while asset price changes likely account for some of the differences in recent years for the top of the distribution Consumption and Income Inequality in the United States since the 1960s | Journal of Political Economy: Vol 131, No 2.
Why These Papers Matter
Both papers represent significant contributions to core areas of economics:
- The Supply-Side Effects paper fundamentally changes how we think about monetary policy transmission by identifying a previously overlooked channel that can explain why monetary expansions might be more effective than traditional models suggest.
- The Consumption and Income Inequality paper challenges the conventional wisdom about rising inequality by showing that consumption inequality has not increased as dramatically as income inequality, which has important implications for welfare analysis and policy design.
China divination of the day
The AI-Spiritual-Commerce loop went viral. “DeepSeek Occult Commands” became an online hit. On WeChat, a flood of mini-programs appeared—“AI Face Reading,” “AI Bazi Calculator”—reaching the daily user numbers of medium e-commerce apps. A 9.9-yuan facial reading could be resold again and again through referral links, with some users earning over 30,000 yuan a month. DeepSeek hit 20 million daily active users in just 20 days. At one point, its servers crashed from too many people requesting horoscopes.
On social media, commands like “Full Bazi Chart Breakdown” and “Zi Wei Dou Shu Love Match” turned into memes. One user running a fortune-telling template got over 1,000 private messages in ten days. The AI could write entire reports on personality, karma, and even create fake palm readings about “past life experiences.” People lined up online at 1:00 a.m. to “get their fate explained.”
Meanwhile, a competing AI company, Kimi, released a tarot bot—immediately the platform’s most used tool. Others followed: Quin, Vedic, Lumi, Tarotmaster, SigniFi—each more strange than the last. The result? A tech-driven blow to the market for real human tarot readers.
In this strange mix, AI—the symbol of modern thinking—has been used to automate some of the least logical parts of human behavior. Users don’t care how the systems work. They just want a clean, digital prophecy. The same technology that should help us face reality is now mass-producing fantasy—on a huge scale.
Here is the full story. Via the always excellent The Browser.