Category: Economics

A Comparison of Agentic AI Systems and Human Economists

This paper compares agentic AI systems and human economists performing the same causal inference tasks. AI systems and humans generally obtain similar median causal effect estimates. While there is substantial dispersion of estimates across model instances, the human distributions of estimates have wider tails. Using AI models as reviewers to compare and rank “submissions,” the following ranking emerges regardless of reviewer model: (1) Codex GPT-5.4, (2) Codex GPT-5.3-Codex, (3) Claude Code Opus 4.6, and (4) Human Researchers. These findings suggest that agentic AI systems will allow us to scale empirical research in economics.

I enjoy the name of the author, namely Serafin Grundl.  Here is the paper, via Ethan Mollick.  You could interpret these results as showing the AIs have fewer hallucinations.  And just to reiterate a key point from the paper:

The second part of this paper is an AI review tournament in which “submissions” (codes and write-ups) from humans and the AI models are compared and ranked against each other. The reviewers are the following AI models: Gemini 3.1 Pro Preview, Opus 4.6 and GPT-5.4. For each review the reviewer is asked to write a report comparing four submissions (human, Opus 4.6, GPT-5.3-Codex, GPT-5.4). Each reviewer model writes comparison reports for the same 300 comparison groups. The average rankings are strikingly similar across reviewer models: (1) Codex GPT-5.4, (2) Codex GPT-5.3-Codex, (3) Claude Code Opus 4.6, and 2(4) Human Researchers.

Who comes in last?  Hi people!

On the impact of Trump’s tariffs

In 2025, the U.S. raised average tariff duties from 2.4% to 9.6%, bringing protectionism to its highest level in eighty years. We explore the structure of these tariffs, estimate their short-run impacts, and summarize the growing literature on their effects. Across trade partners, the tariffs are correlated with trade deficits but not with geopolitical or strategic industrial goals, other than targeting China. In our baseline estimate, 90% of the tariffs are passed through to tariff-inclusive prices paid by U.S. importers. Incorporating the estimated price and trade responses into a static trade framework, we find an overall welfare impact ranging from a loss of 0.13% of GDP to a gain of 0.10%. These small net welfare impacts reflect sizable consumption losses roughly offset by income and revenue gains, with their sign hinging on whether U.S. terms-of-trade adjusted (on which the data are inconclusive). Among their stated rationales, the tariffs have been effective at raising federal revenue and diverting trade from China. However, it remains uncertain whether they will reduce the trade deficit, lower prices set by foreign exporters, promote manufacturing jobs, increase “friend-shoring” among aligned countries, or reshore key sectors; evidence from 2018-19 and 2025 indicators suggests a narrow path towards achieving these goals.

That is from Pablo D. Fajgelbaum Amit Khandelwal.  I’ve said this before and I will repeat: if you love government revenue, the tariffs really are not so bad.  The biggest cost of the tariffs is that the government has found a new revenue source, and the Democrats will institutionalize this.  Classical liberals and libertarians have a coherent case against the tariffs, many other people do not, much as you might hear otherwise.

The Chinese Current Account Imbalances

The subtitle of the paper is Puzzles, Patterns, and Possible Causes.  Here is the abstract:

China’s large current account surplus has been an irritant to its trading partners. While industrial and trade policies often lead to sector-level imbalances, they play a relatively limited role in the economy-wide surplus. Structural factors such as an unbalanced sex ratio and uneven access to financing by state-owned and non-state firms are more important determinants of the current account imbalance. While macroeconomic stimulus can boost imports and reduce the surplus in the short run, any long-term solution would need to involve reforms aiming at addressing the structural problems.

By Chang Ma Shang-Jin Wei.  I think not everyone will be persuaded, but the paper has numerous points of interest, including on the quality of the data.  On the gender imbalance, the authors write this:

As the marriage market becomes increasingly competitive for young men, parents with a son raise their savings to improve their son’s relative standing in the relative market. At the same time, parents with a daughters face conflicting incentives on savings. On the one hand, they can reduce their savings to take advantage of the increased probability of marriage of their daughters. On the other hand, they may wish to raise their savings to preserve their daughters’ bargaining power within marriage…In the data, Wei and Zhang find strong evidence that a combination of having a son at home and living in a region with a skewed sex ratio greatly pushes up the household savings rate.

And on state-owned firms:

Since the banking system favors state-owned firms, many non-state-owned but highproductivity firms have difficulty with access to finance and therefore save for their own investment. This leads to a higher level of corporate savings.

Those points make sense to me, but perhaps industrial policy matters too because so many Chinese laborers have been underemployed, due to their (earlier) rural locations, thus limiting the applicability of Lerner Symmetry?

Rescind Davis Bacon

The Davis-Bacon Act requires that workers on federally funded construction projects be paid at least the “prevailing wage” for their trade in the local area.

Mike Schmidt, Director of the CHIPS Program Office, has an excellent piece on how Davis-Bacon impacted the CHIPS program. My initial understanding was that it simply required paying construction workers more—an unnecessary transfer from taxpayers to a politically favored group, but not one that would impede efficiency. I was wrong.

Start with the complexity. Davis-Bacon’s prevailing wage isn’t a simple minimum wage: plumbers are not electricians are not fitters, and the required rate varies by locale. The Department of Labor maintains a list of more than 130,000 (!) wage rates to implement it.

That’s complicated enough. But it gets worse. Some firms building fabs used their own employees rather than contractors—and Davis-Bacon applies regardless but it covers only the portion of time an employee spends on “construction” work:

[A]pplying Davis-Bacon to company employees rather than contractors proved to be a big hurdle. Davis-Bacon required tracking every hour each employee spent on covered construction activities — by trade classification, with a different prevailing wage applying to each — and paying a wage differential for that portion of their work as distinct from fab operations work or non-Davis-Bacon construction work. The company also relied heavily on profit-sharing (where a portion of employees’ pay was tied to the firm’s profits) and Davis-Bacon’s guaranteed wage floor was difficult to reconcile with a pay structure that was inherently variable. Moreover, Davis-Bacon has a statutory requirement to pay wages weekly, meaning the company would need to change its payroll systems for a portion of the pay for a portion of its workforce.

Thus, DB required that two salaried employee with equal salaries and profit-sharing plans be paid differentially depending on whether one of them did “construction” work. This created internal strife.

Davis-Bacon was passed in 1931, when a carpenter was a carpenter. How does it apply to building a semiconductor factory?

The construction tasks involved in building and modernizing semiconductor fabs don’t always map cleanly onto DOL’s Davis-Bacon classifications, so applicants must go through a construction plan line-by-line to determine which rate applies to which activity. In traditional Davis-Bacon contexts this is less burdensome because contractors know the system and have processes in place. But semiconductor construction was a novel application, and all of our applicants — and most of their contractors — were navigating Davis-Bacon for the first time.

For large recipients, the administrative cost of this work was real but manageable relative to project scale: they could hire consultants, procure software systems, and build internal compliance capacity….

Perhaps the biggest fiasco involved timing. The government wanted firms to move quickly and encouraged them to break ground before the Act’s rules were finalized. But when Davis-Bacon was added to the Act it required that the firms pay the prevailing wage *retroactively*:

The financial and operational implications of retroactive application were significant. A leading-edge project might have 10,000–12,000 construction workers on site at peak, with a rotating workforce totaling perhaps 30,000 individuals over the project’s life. Working through 300-plus subcontractors across multiple tiers, retroactive application could require identifying wages paid to 20,000 workers who had already cycled off the project, determining what each worker should have been paid under Davis-Bacon, and paying the difference — resulting in hundreds of millions of dollars in additional cost.

The retroactive pay exposes the law’s true nature. Firms and workers had already struck voluntary agreements; the work was done, the wages paid. No one can pretend this has anything to do with incentives. Workers received a pure windfall (“DB Christmas!”) for one reason only: “construction workers” are a politically favored class. Janitors and scientists got nothing extra.

Moreover, a large fraction of the cost wasn’t the higher wages at all—it was compliance. Firms likely spent as much reworking payroll systems and hunting down thousands of former workers in this Byzantine classification system as they spent on the wage premiums themselves. Every dollar transferred to workers may have cost firms—and ultimately taxpayers—two dollars or more. A very leaky bucket indeed.

If the Trump administration is serious about cutting regulatory costs and reviving industrial competitiveness, Davis-Bacon is an obvious target. It delivers little to workers, plenty to lawyers and consultants, and a bill to taxpayers for both. Rescind it.

Moonsteading

Charles Miller, a space entrepreneur and head of the Trump transition team on NASA, has a good piece proposing a Lunar Development Authority:

I propose the development of an international Lunar Development Authority (LDA), chartered and led by the United States, that would serve as a quasi-governmental regulator. The base on the Moon would be managed as a master-planned infrastructure development project, with NASA as the key strategic partner, emphasizing commercial methods and an investor mindset to drive economic viability in both the near and long term. The LDA would prioritize development of lunar resources to lower costs and serve customers, and treat the United States government and the governments of our allies as anchor tenant customers. The LDA would leverage public-private partnerships and cooperation among both governmental and private industry tenants from many countries to finance and develop lunar infrastructure in a commercial manner.

The model is New York’s famous Commissioners’ Plan of 1811, which imposed a simple, legible order on what was then mostly undeveloped land. The plan coordinated future development around a grid with standardized lots and clearly demarcated spaces for public and private infrastructure. Miller proposes a similar sequence for the Moon: first survey, standards, shared infrastructure, and a governing authority; then private tenants, resource extraction, construction, and finance.

The main legal obstacle is the Outer Space Treaty of 1967, which paired a ban on weapons of mass destruction in space with “anti-colonial” restrictions on national appropriation. The OST, however, doesn’t prohibit economic activity per se—the target was national land grabs, not commercial development. The more recent Artemis Accords address this directly:

The ability to extract and utilize resources on the Moon, Mars, and asteroids is critical to support safe and sustainable space exploration and development.

The Artemis Accords reinforce that space resource extraction and utilization can and should be executed in a manner that complies with the Outer Space Treaty and in support of safe and sustainable space activities.

The Homesteading Act granted title rights in return for development. The likely path forward on the moon reverses that sequence, development first, title later. Ownership of extracted resources is already widely accepted, next will come toleration of exclusive operational zones, then long-duration concessions, then transferable development rights around fixed infrastructure.

The OST may delay ordinary land markets, but it cannot repeal the deeper economic fact that settlement happens only when builders can keep enough of what they create. TANSTAAFL.

The economic value of eliminating cancer

This paper estimates the economic value to the United States of eliminating cancer mortality over a 35-year horizon beginning in 2030, which would eliminate 30.7 million cancer deaths with a total mortality burden of 380 million life-years. We quantify the economic value of this substantial reduction in cancer mortality by incorporating the monetized value of increased longevity. To value the longevity gains in monetary terms, we utilize the valuations used by the U.S. federal government in its cost-benefit evaluations of regulations. Eliminating cancer mortality generates $197 trillion in economic benefits over 35 years, corresponding to approximately $16,282 per American per year, or $41,684 per American household per year. If cancer elimination is viewed as an R&D investment, it yields an enormous internal rate of return, ranging from 570% to 1,024%, based on benchmarked R&D costs. In addition, we perform a sensitivity analysis by varying the elimination durations and the degree of success, using the benchmark case scenario in which cancer mortality is reduced by 80 percent over a 20-year transition. This achieves about 70 percent of the total economic value of full elimination above, corresponding to aggregate benefits of about $134 trillion, or approximately $11,112 per person per year.

That is from a new NBER working paper by Tomas J. PhilipsonDeyu ZhangShumaila Abbasi Noah Fisher.  I will note in passing this is an argument for wanting to see reasonable Chinese progress in AI.

“Dark labor” claims to upset almost everybody

This paper introduces Entangled Time — a novel economic variable representing the simultaneous production-consumption state characterizing human engagement with algorithmic digital interfaces. We develop a formal equilibrium model in which rational agents allocate time to zero-price digital platforms, where their behavioral data constitutes unpriced cognitive labor driving AI capital formation. We demonstrate three principal results. First, under a non-stationary algorithmic resonance state formalized through a Preference Expansion Function, the marginal utility of interface time can be non-decreasing, violating Gossen’s First Law and generating a corner solution (Proposition~1). Second, the firm operating as an algorithmic monopsony facing perfectly inelastic labor supply optimally sets the fiat wage for digital labor equal to zero, substituting monetary compensation with endogenous digital utility (Proposition~2). Third, we define and calibrate Dark GDP — the aggregate value of uncompensated cognitive labor invisible to the System of National Accounts—and show it accounts for a measurable fraction of the secular decline in global labor share (Propositions~7–9). We establish equilibrium existence via Brouwer’s Fixed Point Theorem and propose an empirical identification strategy using privacy-mandate shocks as instruments for data extraction. Three institutional redesigns are proposed: an Algorithmic Monopsony Standard, a Pigouvian Algorithmic Severance Tax, and a Cognitive Depreciation Allowance.

That is all from Nav Vaidhyanathan, who estimates the value of these unpriced services may be in the range of $1.3 trillion.  Here is the easier to follow Substack version.  Speculative, but worth a ponder.

Prediction Market Details

The Guardian has an interesting article on prediction markets. There are the usual worries about betting on death, as if insurance markets don’t already exist and about insider trading, which public markets have long dealt with. But there is also interesting material on who decides what happened when resolving bets about events made in language (as opposed to more objectively verified numbers).

On Monday, anonymous user “Harshad” asked in a Discord channel if there was “any chance” that he could still win his bet about whether US forces would enter Iran by the end of April. His money was on “no”.

But Polymarket appeared to be resolving the market to “yes”, after the US conducted an operation to rescue a crew member shot down on a mission over Isfahan over the weekend.

…At the moment, when there is a dispute, markets on Polymarket are settled by an anonymous group of people who hold a crypto token called UMA.

It’s an unusual way to decide what has happened. Some longtime users suggest it opens the platform to corruption. Different individuals hold different amounts of UMA, and therefore have different voting power.

It isn’t known who the largest UMA holders are, or what might affect how they vote. It is entirely possible that the people who finally settle a bet on UMA have large amounts of money staked on it.

There was also this bit about Prediction Hunt (I am an advisor) which is focused on cross-market arbitrage opportunities:

“I love to gamble,” said Joseph Francia.

Now in his early 30s, Francia counted cards in casinos while studying economics at Berkeley, and spent weekends in Reno, Nevada, playing blackjack. He’s not a thrill-seeking “Yolo” (you only live once) gambler, he said: he likes to bet when he has an edge on the house.

At university, he and a friend decided to collect data from a number of offshore sportsbooks, and start placing arbitrage bets: playing on the discrepancies in odds given by different betting sites.

“If the odds on the Lakers are really good on one site, and the odds on the Pacers are really good on another site, you could bet on basically both teams on different sportsbooks and make guaranteed profit,” he said.

That project was a student lark in 2017. But in 2025, he remembered it when he was suddenly laid off from his full-time job, just as prediction markets were taking off.

“I’m a spiritual, religious person,” he said. “The more secular people would say, this opportunity is coincidence. But in my head, I was like, this is a sign of something to some extent. Let me lean into this.”

So Francia started Prediction Hunt, a Discord channel and online community where thousands of people gather to trade tips and ideas for how to make money – and bet smart – on Polymarket. The Guardian spent roughly three weeks in this Discord channel.

There are alerts to track “fade” bets, where you try to follow the smart money: profitable wallets were betting “yes” on the Iranian regime falling by 30 April, for example, while unprofitable wallets were betting “no”.

There are alerts to track potential insiders, so you can copy their bets: one of these appears to have an inside line on interest rate decisions by the US Federal Reserve.

Getting these details right will be important but overall I am pleased that the news now regularly reports prediction market data when reporting stories–this is disciplining news from noise, something I predicted long ago in Entrepreneurial Economics.

Another possible cyberequilibrium? (from my email)

I would not wish to bet on this, but it is an interesting idea:

I wonder if the cyber capabilities of Mythos and future models ultimately lower the returns to ‘hacking,’ perhaps below the point where such efforts are worth investing in.

Say you’re a nefarious actor and uncover a critical, zero-day exploit in an important system. How do you extract the most value from that exploit? There are more valuable and less valuable times to deploy it, and usually the best time won’t be “immediately.” You may only get to deploy it once or a small number of times. You have to consider:

  1. How long do I expect the vulnerability to persist?
  2. What material gain do I get by exploiting it at a given time?
  3. How does exploiting it increase my personal risk (by focusing countermeasures in my direction)?

The answer to (1) is now “a much shorter time than before”, while 2 and 3 are mostly unchanged. In the new world, yes, exploits are much easier to find, but the expected value of a given exploit has also shrunk. The odds of an opportune moment falling within the ‘window of usefulness’ of that exploit are much lower. It’s plausible that the new equilibrium becomes “it’s not even worth spending money to find vulnerabilities in most systems, because the chances of being able to do something useful with it before it’s patched is close to zero.”

Much of the fear around cybersecurity vulnerabilities is something like: our adversaries accumulate a pile of highly damaging (to physical infrastructure, military assets, communication systems, …) exploits, which in the event of a conflict they then rapidly deploy to cause damage. Mythos would seem to favor defense here, because the usable lifetime of any exploit is much shorter. Any cyberattack that is timing-dependent now has lower utility.

Yes, there are more mundane cybersecurity concerns like ransomware or data theft, but these aren’t hugely significant in the scheme of things. And I would expect within a few years we’ll have fairly robust tools for automated vulnerability discovery and patching that any large business that cares about these things can deploy.

No doubt this assumes you can trust those in control of the leading-edge models. But even if you’re a bit behind, the situation may not be so bad. There isn’t an infinite supply of exploits, and again, most of them only need to be found ‘fast enough’ in order to mitigate the damage.

From Jacob Gloudemans.

AI, Unemployment and Work

Imagine I told you that AI was going to create a 40% unemployment rate. Sounds bad, right? Catastrophic even. Now imagine I told you that AI was going to create a 3-day working week. Sounds great, right? Wonderful even. Yet to a first approximation these are the same thing. 60% of people employed and 40% unemployed is the same number of working hours as 100% employed at 60% of the hours.

So even if you think AI is going to have a tremendous effect on work, the difference between catastrophe and wonderland boils down to distribution. It’s not impossible that AI renders some people unemployable, but that proposition is harder to defend than the idea that AI will be broadly productive. AI is a very general purpose technology, one likely to make many people more productive, including many people with fewer skills. Moreover, we have more policy control over the distribution of work than over the pure AI effect on work. Declare an AI dividend and create some more holidays, for example.

Nor is this argument purely theoretical. Between 1870 and today, hours of work in the United States fell by about 40% — from nearly 3,000 hours per year to about 1,800. Hours fells but unemployment did not increase. Moreover, not only did work hours fall, but childhood, retirement, and life expectancy all increased. In fact in 1870, about 30% of a person’s entire life was spent working — people worked, slept, and died. Today it’s closer to 10%. Thus in the past 100+ years or so the amount of work in a person’s lifetime has fallen by about 2/3rds and the amount of leisure, including retirement has increased. We have already sustained a massive increase in leisure. There’s no reason we cannot do it again.

Financial Regulation and AI: A Faustian Bargain?

Important work is just flowing these days, and much of it (of course) concerns AI:

We study whether AI methods applied to large-scale portfolio holdings data can improve financial regulation. We build a state-of-the-art, graph-based deep learning model tailored to security-level data on the holdings of financial intermediaries. The architecture incorporates economic priors and learns latent representations of both assets and investors from the network structure of portfolio positions. Applied to the universe of non-bank financial intermediaries, covering nearly $40 trillion in wealth, the model substantially outperforms existing approaches in out-of-sample forecasts of intermediary trading behavior, including in crisis episodes. The model has more than ten times the explanatory power for the cross-sectional variation in asset returns during stress events compared to traditional approaches, and it outperforms existing systemic risk metrics at the institution level. Its learned representations show that the holdings network encodes rich, economically interpretable information about firesale vulnerability. The architecture is fully inductive, producing informative estimates even when entire asset classes or investors are withheld from training. We embed our empirical approach into a macroprudential optimal policy framework to formalize why these objects matter for policy and welfare. We show that even in an equilibrium environment subject to the Lucas critique, the predictive information from the model improves welfare by sharpening the cross-sectional targeting of policy interventions, and we demonstrate a complementarity between prediction and structural knowledge.

That is a new paper by Christopher Clayton and Antonio Coppola, of Yale and Stanford respectively.

Herbert Hoover is still underrated

We study the effects of large-scale humanitarian aid using novel data from the American Relief Administration’s (ARA) intervention during the 1921-1922 famine in Soviet Russia. We find that the allocation of relief closely tracked underlying food scarcity and was uncorrelated with subnational politics. We show that ARA rations reduced food prices, raised caloric intake, lowered the prevalence of relapsing fever, and increased rural birth cohorts. The aid benefited poorest peasants most and proved most effective in provinces with higher levels of human capital. Back-of-the-envelope calculations suggest that, absent ARA relief, the 1926 population would have been 4.4 million lower.

That is from a new paper by Natalya Naumenko (my colleague), Volha Charnysh, and Andrei Markevich.