Category: Web/Tech

How well does AI peer review work?

Claude and I planted 100 known errors into 10 open-access psychology papers and then ran them through frontier models and two commercial AI review tools. In brief:

  • The best single system caught 71 of 100 errors, while the worst caught 30.
  • Pooling every system’s output caught 93 of 100. Models are only partly correlated in the errors they find, making ensembling a big lever for finding issues in papers. Check your papers against multiple models!
  • Seven errors could not be caught by any system. All were omissions — information deleted from a paper rather than mistakes inserted into it.
  • Refine.ink contributes more unique catches than any other single system, though it’s expensive.
  • I didn’t measure false positives and I don’t know how this error distribution compares to the distribution of errors in real papers.
  • I’ve made the papers, errors, model outputs, and the full experiment log public. I hope people can build on this work to create a comprehensive eval benchmark across disciplines.

That is from Paul Litvak, here is more.  Note that is not even using the very latest generation of models.

Act like it is science

I am seeing so many doom prognostications, or least severe worries, due to the OAI/HuggingFace incident and related stories.  But virtually all of these I find underargued to say the least.

So I have a simple request.  If you are worried, and wish to persuade the doubters, try the methods of science.  I would like to see the following:

1. Your outline of how, say two or three years from now, we might estimate the additional cybersecurity costs from the AI break-ins.  I am convinced that number is not zero, but give me your method please.  If it helps, here is an estimate of past cybersecurity costs, done by top economists.

2. Your current numerical estimate of what those costs might end up being, of course to be tested against what actually happens over time.  Obviously, you can do this for a few different regulatory/safety scenarios.  This is one simple way to prove yourself largely correct, albeit with a lag.  (NB: you do not have to take this as a substitute for your preferred safety measurres.  But please try to be specific in your predictions.)

3. A list of what stocks or other assets you have shorted, now.  Obviously if your answer to #2 is sufficiently low, you could answer here zero, as I would do.  I expect costs, but not so high that we cannot muddle through and have expected positive stock returns.

This would all help to make the discourse more scientific and less like an extended exercise in personal anxiety management.

If it is all so important, surely this exercise is worth the effort?

The future of warfare? (from my email)

As weapon systems increase in effectiveness, they will generally target one another rather than humans. Lower loss of life will ease conflict entry and hinder exit.

Warfare will veer further from industrial competition into techno-industrial; the current appearance of parity between low tech and advanced actors reflects near zero deployment of advanced systems. As these deploy, their martial advantages will present clearly.

The need for advanced ecosystems will lead to blocs able to combine human capital, techno-industrial capacity, and inputs. Logistics may be contested at elevated levels, changing the incentives around geographic size and increasing the value of inputs.

Blocs may well engage in near constant warfare of various intensity levels. Bloc composition will be necessity driven, perhaps with increased fluidity.

That is from an anonymous reader.

*Restoring Childhood: How to Set Kids Free in the Age of Anxiety*

That is the new book by Peter Gray, and it is very much a whole philosophy of child-rearing and also antidote to the current panic over social media. Excerpt:

In preparation for this chapter, I spent weeks poring over the research literature pretaining to smartphones, social media, and teen mental health.  It is clear to me that there is something close to a consensus among scientists most immersed in this research, which includes the following three conclusions: (1) Extensive research refutes the theory that effects of smartphones and social media account for a meaningful portion of the decline in teen mental health in the years following 2010.  (2) Smartphones and social media can have both beneficial and harmful effects on teens’ mental health, which tend to cancel out when averaging.  (3) Future research should focus on understanding better the ways teens use these tools and how to help them maximize the benefits and minimize the harm.

Note that teen suffering also hit a peak around 1990 by many measures, yet that cannot have been caused by phones or social media.  And European teen suicide rates were falling just as those in the United States were rising.  This evidence is considered in further detail of course.

Recommended, it is both easy to read and substantive.

Mexico (Taiwan) fact of the day

Mexico has quietly become a cornerstone of the AI boom, providing 40 per cent of US imports this year of the computer servers that are widely used in the data centres powering artificial intelligence.

Taiwanese manufacturers are rapidly expanding factories in Mexico to assemble servers, which are now the country’s top US export, overtaking autos that dominated its trade for decades.

Mexico is the second-largest provider of enterprise servers to the US so far this year, with sales reaching $46.9bn, behind Taiwan, which sold $53.5bn, though Mexico became the largest on a monthly basis in May.

The phenomenon is helping prop up Mexico’s limp economy, pushing the country’s exports to record levels. Servers and related hardware made up almost one-fifth of the $317bn of goods Mexico exported between January and May, more than double the same period a year earlier. Taiwan, in turn, is now Mexico’s third-largest trading partner, up from eighth place in 2022. Taiwanese companies have spent more than $1.6bn since 2020 on Mexican factories that offer geographic proximity and tariff-free access to US tech giants spending hundreds of billions on data centres.

“Mexico is of tremendous importance to the way this new AI-powered economy is working,” said Jesse Rogers, head of Latin America economics at Moody’s Analytics. “It is really a new frontier of collaboration and even dependency on the Mexican economy when it comes to AI servers.”

Here is more from Christine Murray and Alan Smith at the FT.

The wisdom of Eli Dourado

How do “international efforts” work? Unlike most of the signatories of the letter, I have been a State Department advisor and have participated in multiple treaty negotiations. I have also been a member of technical committees for UN technical agencies.

The international sector is unbelievably dysfunctional. Every single treaty or international agreement is an opportunity for every participant to manipulate the much broader policy environment. Often the participants don’t care about the object of the agreement, and are instead trying to use the agreement as leverage for something else. I have seen countries use their limited leverage in multilateral agreements to try to kneecap American industry, get around sanctions, create a pretense of international justification for domestic illiberalism, or steer technology in an authoritarian direction.

The damage from this dysfunction is limited by the fact that there are zero major countries (and not that many smaller countries) who will actually bind themselves in meaningful ways that they don’t narrowly want. If a previously signed treaty turns out to be inconvenient, it is often subtly ignored or reinterpreted.

In addition, treaty delegations working on industrial issues generally reflect the full spectrum of special interests involved in an issue. A US AI treaty delegation might be led by a State Department ambassador, but he will be advised by representatives of the interagency, major labs, other major tech companies, tech investors, civil society, etc.

“International efforts,” therefore, is not a reassuring answer to the governance problem; it is the name of another enormous, unresolved governance problem.

Here is the full tweet.

Data on Chinese innovation

China’s technological progress in recent decades has been viewed with admiration, alarm, and (in some cases) doubt. To better understand the Chinese innovation ecosystem, we compile a dataset of almost 14 million domestic Chinese patent publications. We focus on the subset of critical technologies identified by the U.S. Department of Defense. Several surprising patterns emerge from the data: Chinese patenting is strongly associated with other measures of innovative progress; patents are not concentrated in corporate giants such as Huawei; universities have played a key role in innovation, much greater than state-owned enterprises or government-owned facilities; and fewer than one in ten Chinese critical technology patents involves an inventor with U.S. experience or training. Finally, using four text-based measures of patent quality, we show that the rise of Chinese patenting in critical technologies has not been associated with a decline in quality relative to the U.S. awards.

That is from a recent paper by Josh Lerner, Namrate Narain, Dimitris Papanikolaou, Amit Seru, and Zunda Winston Xu.  Via the excellent Kevin Lewis.

You will learn to love AI writing

That is the theme of my latest Free Press article, excerpt:

But do I wish to eschew AI writing for the rest of my life? Absolutely not. Most of all, I want AI writing to get better, so it does not irritate me with its clichés and all too obvious identifying marks. I want AI writing that can fool me, and maybe sometimes I am already getting it.

Do you know Paul Simon’s lovely song “American Tune”? It is one of my favorites. A small percentage of listeners know that one of the core melodies is taken from Johann Sebastian Bach, namely the Passion Chorale from “St. Matthew Passion.” You could say Paul Simon is a plagiarist, though I would prefer not to use this term. He borrowed creatively, as do many other popular music artists.

Do many listeners care? No, although a few might feel a pang of musical nerd pride at learning of the association with Bach.

But wait, it gets worse yet. That melody does not come originally from Bach, but rather he in turn plagiarized a 1601 secular love song by Hans Leo Hassler, namely “Mein G’müt ist mir verwirret” (My feelings are confused). So the lineage is not even the glamorous one you might have imagined. If anything, “American Tune” sounds closer to the Hassler melody than to Bach’s adaptation.

Are you upset yet? Probably not, and that is good.

Someday we will feel the same way about text generated with the assistance of AI. We won’t care much where it came from, and much of the time we will not even know.

Recommended, and I hope the AIs like it too.

The demand for human enhancement technologies

When a new technology promises large private benefits but may impose social costs that markets do not price, demand need not reveal how citizens want it governed. We examine this using a nationally representative U.S. survey experiment (N=5,556) on human enhancement technologies (HET). The experiment randomizes benefit domain, mechanism, heritability, purpose, and risk across vignettes; for each respondent’s assigned vignette, we elicit stated adoption, preferred regulation, and ethical and societal concerns. Overall, about 53% would adopt. Framing the technology as enhancing rather than restorative lowers adoption by about five percentage points, as much as a severe side-effect profile. About 28% would not adopt at any benefit. This refusal is driven overwhelmingly by the enhancing framing rather than by risk, consistent with a non-compensatory constraint for a substantial subgroup. Most who would adopt still favor strict regulation, and most who would never adopt do not wish to forbid others from doing so. Productivity enhancement generates the most ethical concern of any attribute but attracts the least regulation, and respondents favor subsidizing rather than taxing its adoption, consistent with a concern about access rather than safety. Private demand is therefore an unreliable guide to the governance citizens want, and the divergence we document provides a basis for regulators seeking to align the direction of technical change with societal values and priorities.

That is from a new NBER working paper by Giovanni ImmordinoMario MacisImmacolata Marino Fabrizio Panebianco.  And I will repeat this segment: “Productivity enhancement generates the most ethical concern of any attribute…”  Do note of course that the last sentence of the authors is completely unwarranted, and is a classic example of underidentified political bias in academic reasoning.

The common sense of Fareed Zakaria

This is one of the best Op-Eds you will read this year, though it deliberately deemphasizes the personal (no one is attacked) in a manner that will result in less attention.  Better to write the truth!  Excerpt:

Why has this [voter discontent} happened? Some of it is economic. Expensive housing, stagnant incomes and widening inequality have fueled public anger. But economics is only part of the story. America has enjoyed stronger growth than most other advanced economies. France under President Emmanuel Macron has lowered unemployment, and it has attracted more foreign investment projects than any country in Europe for seven straight years, according to an Ernst & Young analysis. South Korea is an economic success story. Japan was experiencing a comeback. Yet democratic governments almost everywhere are unpopular.

The larger explanation is that we are living through one of history’s great technological revolutions. And these transformations always produce deep anxiety and fear.

The last technological transformation on this scale came between roughly 1880 and 1920. Electricity, the telephone, railroads, the automobile and movies overturned entire industries and ways of life. Millions left farms for cities. Global trade disrupted local economies. New fortunes appeared overnight, while old occupations disappeared. In this turmoil, frustration and anger found political pathways. The left turned to communism, the right to fascism. What followed were new mass movements, cultural upheavals and world war.

Do read the whole thing.  You can be a victim of these processes, or keep your wits about you.  The choice is yours.

The optimal Bayesian update?

I see at least three updates one might make from the recent OAI/Hugging Face hacking incident:

1. “This happened sooner than I expected, and the story is more dramatic than I expected,” therefore I am more worried than before.

2. “This happened, and the inferior Chinese cyber-defense seems to have performed just fine,” therefore I am less worried than before.

3. “This happened, and as far as we can tell, absolutely no one was harmed,” therefore I am less worried than before.

Obviously the net impact, from those bare hypotheses, is indeterminate.  And yet few people seem to be paying much heed to #2 or #3.  Joshua Saxe from the cyber world has some relevant observations.  And perhaps other updates are needed as well.