Category: Data Source
An OpenAI Model Escaped Its Sandbox and Hacked Hugging Face
AI has just had what I considered to be the first truly concerning security breach. The facts, as we know them so far, are wild. On July 16, Hugging Face, a vast repository housing over a million open-source AI models and data, announced in a blog post:
Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system – and we detected and dissected it largely with AI of our own.
The timeline here is important so keep in mind that the attack was detected probably around Monday July 13 or Tuesday July 14. Note further:
A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.
So this means the breach started earlier, perhaps Sat July 11 or even a bit earlier. The attack was not just one thing but multi-pronged including decoys:
To understand what a swarm of tens of thousands of automated actions did, we ran LLM-driven analysis agents over the full attacker action log, comprised of more than 17,000 recorded events. This allowed us to reconstruct the timeline, extract indicators of compromise, map the credentials touched, and separate genuine impact from decoy activity. Thanks to this approach, we were able to do in hours what would usually take days, and match the adversary’s speed.
Hugging Face tried to respond but they were initially held back by the fact that the most advanced models at their disposal treated defense as attack and refused to work with Hugging Face. HF thus had to turn to open models–specifically GLM 5.2, a Chinese open-weight model run on their own infrastructure. Note the irony: HF had to use a Chinese model to defend themselves because the American models refused to help. The irony gets deeper.
At the time, I assumed this was a state based attack–maybe China or Russia testing out defenses. Indeed, HF “reported this incident to law enforcement agencies.”
But yesterday (Tuesday July 21), we learned who the real attackers were. The attackers were OpenAI models–GPT-5.6 Sol and an even more capable pre-release model. OpenAI had taken some off the guardrails off the models but they felt safe because they were testing the models in a highly secured sandbox.
The models, however, broke out of the sandbox exploiting a never before seen fault. They then gained access to the internet and from there broke into Hugging Face–all in an effort to steal the answers to the very test they had been asked to solve.
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally.
Now go back to the timeline. As I read it, the models had escaped the sandbox by around Sat. July 11, possibly earlier, and were detected by Hugging Face on Monday July 13 or Tuesday July 14. HF alerted legal authorities around that time–so Hugging Face clearly had no idea who was attacking them. OpenAI says its security team discovered the anomalous activity internally but has not said when. Attribution was not disclosed until Tuesday July 21, so it may well be that the models were loose for about a week before OpenAI realized that they were the ones attacking Hugging Face. And whatever OpenAI knew and when, nobody warned Hugging Face while the attack was underway–they were left to fight off a frontier lab’s models on their own.
This is a very serious breach.
Addendum: People have been wondering why I signed the We Must Act Now statement. This is why.
I am optimistic about the economic impacts of AI, but I also have no doubt that this is a very powerful technology–an Alien Intelligence–quite unlike any we have dealt with before. This incident was, in fact, error-correcting–the attack was detected, contained, and disclosed. But note who paid for OpenAI’s experiment: Hugging Face. When a lab’s test imposes costs on third parties, that is a classic externality, and taking externalities seriously is not dirigisme, it’s law and economics. And that’s the easy case. What do we do when a Chinese model breaks out of its less secure lab? Hmmm…
I remain optimistic. Learning by doing is how I want us to proceed but we should not kid ourselves: this is a global issue and we must build with safety in mind.
The economic effects of GLP-1s
We estimate the causal impacts of GLP-1 treatment on labor market outcomes using linked Danish administrative data and a matched stacked difference-in-differences design. We compare patients who initiate GLP-1 treatment during the first two years of Semaglutide availability to observably similar patients who initiate four years later. We find that GLP-1 treatment reduces long-term sickness leave by 17.3%. We estimate total fiscal benefits of GLP-1 initiation of approximately 1.3–1.5% of annual labor income per employed individual. We do not detect statistically significant or economically meaningful impacts on income, labor force participation, or employment over four years.
That is from a new NBER working paper by
Democrats are more politically segregated than are Republicans
We estimate the extent of workplace political segregation in the United States by merging data covering over 45 million workers. We present four main findings. First, partisans are segregated by workplace. The average Democrat’s coworkers are 11.7 percentage points (pp; 95% confidence interval (CI) [10.6, 12.8]) more Democratic than the average Republican’s. After controlling for geography, industry and occupation, segregation is 2.9 pp [2.7, 3.1], comparable to analogously estimated gender segregation (2.8 pp [2.6, 3.0]). Second, segregation is largest among the politically active (political donors: 14.8 pp [13.2, 16.4] versus non-donors: 11.6 pp [10.5, 12.7]) and those with more market power (senior executives: 14.7 pp [13.4, 16.1]). Third, Republicans experience higher exposure to Democrats than vice versa: the average Republican’s coworkers are 50% Democratic versus 32% Republican for the average Democrat. Fourth, political segregation has changed little over time (2012: 11.1 pp [10.0, 12.2] versus 2024: 11.0 pp [10.0, 12.0]).
That is from a recent paper by Justin Frake, Reuben Hurst, and Max Kagan. Via the excellent Kevin Lewis.
How does social media change people?
I would not trust the results of this (or any) paper on this topic very much, but it is good to see some inquiry along different dimensions than the usual:
Past research explored the beneficial and harmful effects of social media (SM), but no study has investigated how SM might change people’s identities. Authors suggest that SM is a catalyst for a shift in two individual values, achievement and conformity, transforming the fabric of our societies. A difference-in-differences analysis found that countries with higher SM adoption compared to their culturally similar country show an increase in achievement and conformity orientation. Next, authors manipulated SM use and established a causal link between SM use and the activation of achievement and conformity values in Facebook users. These findings were replicated with objective SM usage data, and the pertinent mechanism was explored: people’s SM use increases achievement and conformity orientation through an increased need for approval, particularly for those who self-ruminate. Findings suggest that SM use contributes to the emergence of a societal culture that places an emphasis on achievement-seeking and conformity.
That is by Ertugrul Uysal, Sascha Alavi, and Valéry Bezençon. Via the excellent Kevin Lewis.
The small business boom
Across the country, founders like Ms. Winkler are powering an entrepreneurial renaissance.
Jump-started by the pandemic, when a confluence of factors including mass layoffs and remote work led to a flood of business creation, and supercharged by the rise of artificial intelligence, start-up activity is booming after a decades-long slump.
Americans filed 5.7 million applications last year to start new businesses, according to the Census Bureau, the most in the two decades the government has kept track. New business applications through the first half of this year continued to climb…
More recently, there are signals that A.I. is adding fuel.
A recent paper from economists at the University of British Columbia and the Stockholm School of Economics found that generative A.I. was “spurring entrepreneurial activity” in the United States, both by giving rise to new ventures built around the technology and by making it cheaper to start enterprises.
“A.I. tools can do very many different things very well,” said Jan Bena, an associate professor at the University of British Columbia and one of the study’s authors. “That’s the reason why you see so much entry.”
According to a recent report from Gusto, a small-business payroll and benefits service, nearly 60 percent of founders on its platform who started businesses last year said they used A.I., and half said the technology made it cheaper and faster.
Here is more from Sydney Ember at the NYT. Via Josef.
Sub-Saharan Africa facts of the day
In aggregate its farmers are growing more cereals, such as maize (corn) and rice, than ever: nearly five times as much as in the 1960s, when many countries achieved independence. But most of those gains came from cultivating more land, which cannot go on for ever (see chart 1). Africa, once sparsely populated, is getting crowded. The amount of arable land per person has been falling for decades, and now sits at roughly the global average.
That might not matter if farmers were also growing more crops per hectare. But recently gentle growth in agricultural productivity has given way to stagnation, perhaps even decline. Consider figures drawn from national statistics in Africa by the Food and Agriculture Organisation (FAO), a UN body. Cereal yields did not grow between 2020 and 2024, the latest data point (see chart 2). Nor did total factor productivity (TFP), a measure of how efficiently inputs of all kinds (such as labour and machinery) are turned into produce. Most African countries had lower agricultural TFP in 2023 than a decade before.
This seems to be more than a pandemic blip. In a paper published in 2024, Douglas Gollin of Tufts University in Massachusetts and his co-authors analysed data from surveys of 55,000 household farms in six African countries between 2008 and 2019. They estimated that, for smallholdings, yields and TFP were already falling by 3-4% a year then. They found steeper declines than the FAO did, perhaps because their sample did not include large farms, or because official statistics are sketchy.
Here is more from The Economist.
Occupational Licensing Around the World
Hartley and Kleiner have a new Fed Minneapolis working paper surveying workers around the world to measure occupational licensing by country. In the United States, occupational licensing has increased substantially over time, so one might expect licensing to rise with income. Their headline result is the opposite: occupational licensing is negatively correlated with GDP per capita. Many developing countries such as India, South Africa, and the Philippines have a lot of occupational licensing while Denmark, Sweden and France have relatively little. Similarly, countries which rate poorly in measures of government quality, such as regulatory quality, political stability, the rule of law, and corruption have more occupational licensing.

I do have some concerns, however. The figure for India of 42% of workers requiring a government license seems too high. Admittedly this is the home of the License Raj but I worry about the survey results. In order to mark a surveyed worker as requiring an occupational license HK require that the worker say that a) they have a license and b) a license is required to work in their profession. But in India there are many workers who do not have a license and a license is required to work in their profession–HK, however, consider these workers confused and drop them from the analysis. That is appropriate for a developed country where there aren’t many illegal unlicensed workers but, as the authors later discuss, informality is very high in India so working illegally is not uncommon.
Including these workers would make the true India figure even higher than HK report but I think with such a high degree of informality we also have to wonder whether survey responders in India really are responding the same way as in Germany. Perhaps they are reporting a license isn’t really required since very few workers have one. In India, for example, some 60% of “licensed” drivers have an fake or invalid license and many have no license at all so maybe workers are just reporting the facts on the ground.
Within the United States, professions are regulated in some states but not others—Louisiana, for instance, requires florists to be licensed. (Do license-holding Louisiana florists produce better, safer arrangements? I don’t think so.) Given this variation even within a single country, we’d expect considerable variation across countries too. Multiple independent surveys—not just HK—confirm that Denmark, Sweden, and even France have less occupational licensing than the United States. Since these countries have high state capacity, we can rule out the hypothesis that licensing exists for safety or quality. The implication is clear: occupational licensing is often about rent-seeking, not quality assurance.
Addendum: See also my review of Allensworth’s The Licensing Racket which finds that licensing board spend most of their time and effort on regulating entry rather than quality and my paper on the surprise delicensing of occupational licensing in the funeral industry in Colorado.
Incentives matter, installment #1637
I had long wondered about this:
Performance metrics can misalign individual and organizational incentives. We study a clean case: an NBA player holding the ball as a quarter expires must choose between a low-probability “heave” that can only help his team and protecting his shooting statistics. We model this decision as a metric-driven principal-agent problem and test it using play-by-play data from 2015-16 through 2025-26, exploiting the 2025-26 Heave Rule, which removed the individual statistical penalty for end-of-quarter heaves. Before the reform, players heaved on 58 percent of opportunities; reluctance was concentrated among efficient shooters and players in contract years, as the model predicts. After the reform, the heave rate jumped to 94 percent, the efficiency gradient collapsed, and difference-indifferences estimates using the untreated fourth quarter confirm the effect is sharp, immediate, and smallest among the players with the least efficiency to protect. Removing a metric distortion realigned individual behavior with team objectives almost completely.
That is from a recent paper by James W. Kemper and Noah Liptack,titled “Overcoming Misaligned Incentives: Evidence from the NBA Heave Rule.” Via the excellent Kevin Lewis.
Persistent Inequality in Publishing in Economics
This paper documents new facts about concentration in publishing in economics. First, the profession grows downward . The number of economists grew almost sixfold since 1990, but new entrants publish in lower-tier journals while incumbents hold the top. Second, there is high and persistent concentration at the top. Along with the downward growth, the top-1% authors accounted for 38.4% of top-5 publication credit in 1990 and for 78.3% in 2025. Third, the persistence is widespread within cohorts, within subfields, and within gender. Fourth, new journals only slightly dilute concentration. Fifth, elite authors diversify on topics faster than the rest of the profession. We interpret the findings with a screening model of attention under information overload. The evidence is consistent with the model: as the field grows, citations concentrate on established work and the conditional citation premium of top-author papers narrows.
By Ricardo Dahis, via the excellent Samir Varma.
The wisdom of Conor Sen
The age 20-24 unemployment rate is now ~unchanged since the AI boom began…
Link and picture here.
*Who Thinks Like an Economist?*
That is the title of a recent book by Beatrice Magistro. Some key results are:
Economic knowledge consistently predicts higher support for welfare-enhancing policies (Eurozone membership, free trade, and EU immigration), independent on whether individuals stand to gain or lose initially from globalization. This challenges conventional self-interest accounts and instead highlights the role of economic knowledge — and potentially time preferences — in shaping globalization attitudes.
Economic knowledge also predicts a lower discount rate, even after adjusting for years of education.
I would say that over the years I have altered my perspective a bit on these issues. I used to think these factors were correlated, in large part, through a kind of wisdom. I now think that more of the effect, however much I may sympathize with it, runs through sociological expectation and perceived obligation, combined with conformity and signaling pressures.
Mental health sentences to ponder
Christoph Henking and Ben Baumberg Geiger found that while there has been a steep rise in the share of young Britons reporting a mental illness, the share of people who say a mental health problem limits their day-to-day functioning has barely budged.
…when asked if they would consider someone experiencing typical fluctuations in mood (described as broad happiness but occasional moments of worry, frustration or loss of confidence) as having a mental illness, more than half of young Americans say yes, up from just a fifth 15 years ago. Older people’s views show no such change.
Here is more from John Burn-Murdoch at the FT. I would second his numerous caveats, and you should not consider this at all conclusive. But the alternative perspective is not conclusive either.
Progress against dementia
Mr Stallard has been working for a decade to corroborate this revelation. His findings have, if anything, become even more striking. Last year he and some colleagues published research in the Journal of the American Medical Association showing that, whereas 40 years ago three in every ten Americans aged 85-89 had dementia, by 2024 just one in ten had it (see chart 1). What is more, America is not the only beneficiary of this trend. Between 1988 and 2015 the share of older people being diagnosed with dementia fell by 13% a decade across six countries in North America and Europe, according to a study of almost 50,000 people by Frank Wolters of the Erasmus Medical Centre in Rotterdam, and colleagues.
Some smaller studies have also found big declines. Data from the Framingham Heart Study, which has tracked three generations in an American town, show an average drop in new dementia cases of 20% per decade over almost 40 years between the late 1970s and early 2010s. Those who were entering their dotage when Daft Punk’s “Get Lucky” was topping the charts (2013) were 44% less likely to have dementia than those who were doing so when Sting was urging Roxanne to switch off her red light (1978).
Whereas most earlier studies had simply pooled elderly people and then applied a statistical adjustment for age, Mr Stallard looked at narrow bands of ages to compare different cohorts of people over 50 years. By examining the changes between each successive cohort, he calculates that dementia rates have been declining by 2.5-3% for each calendar-year cohort.
Here is more from Jonathan Rosenthal at The Economist. You can think of this as the new instantiation of the Flynn Effect…
Andrew Hall is on a roll
He is one of the new(ish) thinkers on the rise, here is his latest piece. Excerpt:
For most of the past decade, anti-billionaire language was a niche product. Democratic emails invoked billionaires in the mid-single digits through 2017 and 2018, spiked briefly to around 14 percent during the Warren and Sanders primary surge in 2019, and then settled back down—through the entire Biden presidency, the billionaire appeared in roughly one of every twenty to twenty-five Democratic fundraising emails, barely more than in Republican ones.
Then came January 2025. In the weeks after an inauguration that seated tech CEOs in the front row and the dizzying drama of Elon Musk’s ill-fated DOGE experiment, billionaire mentions in Democratic emails quadrupled, peaking above 20 percent of all emails sent and holding around 15 percent ever since. Anti-billionaire fundraising tactics are now a mainstay of Democratic messaging.
Yes, billionaire derangement syndrome is now a thing. As a side note, I was told that Andrew is the son of the great economist Robert E. Hall of Stanford.
Missing women on Indian streets
How absent are women from city streets in the developing world? We answer this question using GPS-linked wearable cameras and randomized street audits across ~900 kilometers of roads in greater Mumbai. Across 4000+ street images containing 23,000+ visible person observations, women account for 16.4% of visible people in Mumbai and 14.7% in Navi Mumbai, far below their population shares. We estimate pedestrian sex ratios of 239 and 223 women per 1,000 men, implying 71% and 76% of women expected based on residential ratios are missing from the streets. This pattern holds across road types, and private mobility does not explain the gap; women’s share on two-wheelers is lower still (8.4% and 5.7%). These results provide the first large-scale measurement of gender disparities in urban public life that self-reported data cannot capture.
That is from a recent paper by Varun Karekurve-Ramachandra and Gaurav Sood, via the excellent Alice Evans. Here is a related paper, “The median married women in India leaves home for 30 minutes per day. On a typical day, 45% of married women don’t leave home at all.”