The Nuclear Non-proliferation Treaty and existential AGI risk

The Nuclear Non-Proliferation Treaty, activated in 1970, has been relatively successful in limiting nuclear proliferation.  When it comes to nuclear weapons, it is hard to find good news, but the treaty has acted as one deterrent of many to nation-states acquiring nuclear arms.  Of course the treaty works, in large part, because the United States (working with allies) has lots of nuclear weapons, a powerful non-nuclear military, de facto control of SWIFT, and so on.  We strongly encourage nations not to go acquiring nuclear weapons — just look at the current sanctions on Iran, noting the policy does not always succeed.

One approach to AI risk is to treat it like nuclear weapons and also their delivery systems.  Let the United States get a lead, and then hope the U.S. can (in conjunction with others) enforce “OK enough” norms on the rest of the world.

Another approach to AI risk is to try to enforce a collusive agreement amongst all nations not to proceed with AI development, at least along certain dimensions, or perhaps altogether.

The first of these two options seems obviously better to me.  But I am not here to argue that point, at least not today.  Conditional on accepting the superiority of the first approach, all the arguments for AI safety are arguments for AI continuationism.  (And no, this doesn’t mean building a nuclear submarine without securing the hatch doors.)  At least for the United States.  In fact I do support a six-month AI pause — for China.  Yemen too.

It is a common mode of presentation in AGI circles to present wordy, swirling tomes of multiple concerns about AI risk.  If some outside party cannot sufficiently assuage all of those concerns, the writer is left with the intuition that so much is at stake, indeed the very survival of the world, and so we need to “play it safe,” and thus they are lead to measures such as AI pauses and moratoriums.

But that is a non sequitur.  The stronger the safety concerns, the stronger the arguments for the “America First” approach.  Because that is the better way of managing the risk.  Or if somehow you think it is not, that is the main argument you must make and persuade us of.

(Scott Alexander has a new post “Most technologies aren’t races,” but he doesn’t either choose one of the two approaches listed above, nor does he outline a third alternative.  Fine if you don’t want to call them “races,” you still have to choose.  As a side point, once you consider delivery systems, nuclear weapons are less of a yes/no thing than he suggests.  And this postulated take is a view that nobody holds, nor did we practice it with nuclear weapons: “But also, we can’t worry about alignment, because that would be an unacceptable delay when we need to “win” the AI “race”.”  On the terminology, Rohit is on target.  Furthermore, good points from Erusian.  And this claim of Scott’s shows how far apart we are in how we consider institutional and also physical and experimental constraints: “In a fast takeoff, it could be that you go to sleep with China six months ahead of the US, and wake up the next morning with China having fusion, nanotech, and starships.”)

Addendum:

As a side note, if the real issue in the safety debate is “America First” vs. “collusive international agreement to halt development,” who are the actual experts?  It is not in general “the AI experts,” rather it is people with experience in and study of:

1. Game theory and collective action

2. International agreements and international relations

3. National security issues and understanding of how government works

4. History, and so on.

There is a striking tendency, amongst AI experts, EA types, AGI writers, and “rationalists” to think they are the experts in this debate.  But they are only on some issues, and many of those issues (“new technologies can be quite risky”) are not so contested. And because these individuals do not frame the problem properly, they are doing relatively little to consult what the actual “all things considered” experts think.

*A New History of Greek Mathematics*

I have read only about 30 pp. so far, but this is clearly one of the best science books I have read, ever.  It is clear, always to the point, conceptual, connects advances in math to the broader history, explains the math, and full of interesting detail.  By Reviel Netz.  Here is a brief excerpt:

And this is how mathematics first emerges in the historical record: the simple, clever games accompanying the education of bureaucrats.

One of the best books of the year, highly recommended.

Thursday assorted links

1. “We find that using algorithmic responses changes language and social relationships. More specifically, it increases communication speed, use of positive emotional language, and conversation partners evaluate each other as closer and more cooperative. However, consistent with common assumptions about the adverse effects of AI, people are evaluated more negatively if they are suspected to be using algorithmic responses.”  Link here.

2. Living with non-alignment.  And a very sane take on AI risk.  Very good thread.  And should American VCs be funding Chinese AI?  And Leopold Aschenbrenner responds to me on AI risk, very good piece.

3. Why isn’t Europe doing worse?

4. “I hereby challenge professors from universities around the world to submit assignments that they believe are AI-immune.”

5. Which individuals are most likely to believe that AI is likely to destroy society?

Is the Great Awokening a global phenomenon?

And perhaps it did not start in the United States?  Here is more from David Rozado, including a full research paper:

My excellent Conversation with Jessica Wade

Here is the audio, video, and transcript.  Here is part of the summary:

She joined Tyler to discuss if there are any useful gender stereotypes in science, distinguishing between productive and unproductive ways to encourage women in science, whether science Twitter is biased toward men, how AI will affect gender participation gaps, how Wikipedia should be improved, how she judges the effectiveness of her Wikipedia articles, how she’d improve science funding, her work on chiral materials and its near-term applications, whether writing a kid’s science book should be rewarded in academia, what she learned spending a year studying art in Florence, what she’ll do next, and more.

Here is the opening bit:

COWEN: Let’s start with women in science. We will get to your research, but your writings — why is it that women in history were so successful in astronomy so early on, compared to other fields?

WADE: Oh, that’s such a hard question [laughs] and a fascinating one. When you look back at who was allowed to be a scientist in the past, at which type of woman was allowed to be a scientist, you were probably quite wealthy, and you either had a husband who was a scientist or a father who was a scientist. And you were probably allowed to interact with science at home, potentially in things like polishing the lenses that you might use on a telescope, or something like that.

Caroline Herschel was quite big on polishing the lenses that Herschel used to go out and look at and identify comets, and was so successful in identifying these comets that she wanted to publish herself and really struggled, as a woman, to be allowed to do that at the end of the 1800s, beginning of the 1900s. I think, actually, it was just that possibility to be able to access and do that science from home, to be able to set up in your beautiful dark-sky environment without the bright lights of a city and do it alongside your quite successful husband or father.

After astronomy, women got quite big in crystallography. There were a few absolutely incredible women crystallographers throughout the 1900s. Dorothy Hodgkin, Kathleen Lonsdale, Rosalind Franklin — people who really made that science possible. That was because they were provided entry into that, and the way that they were taught at school facilitated doing that kind of research. I find it fascinating they were allowed, but if only we’d had more, you could imagine what could have happened.

COWEN: So, household production you think is the key variable, plus the ability to be helped or trained by a father or husband?

The discussion of chirality and her science work is very interesting, though hard to summarize.  I very much like this part, when I asked her about her most successful unusual work habit:

But just writing the [Wikipedia] biography of the person I was going to work with meant that I was really prepped for going. And if I’m about to see someone speak, writing their biography before means I get this. That’s definitely my best work habit — write the Wikipedia page of what it is that you are working on.

I don’t agree with her on the environment/genes issue, but overall a very good CWT, with multiple distinct parts.

Do women disagree less in science?

This paper examines the authorship of post-publication criticisms in the scientific literature, with a focus on gender differences. Bibliometrics from journals in the natural and social sciences show that comments that criticize or correct a published study are 20-40% less likely than regular papers to have a female author. In preprints in the life sciences, prior to peer review, women are missing by 20-40% in failed replications compared to regular papers, but are not missing in successful replications. In an experiment, I then find large gender differences in willingness to point out and penalize a mistake in someone’s work.

That is from a new paper by David Klinowski.  Via the excellent Kevin Lewis.

Wednesday assorted links

1. Some GPT advances so far.  And good survey of views on existential risk.

2. The paradox of social capital in Northern Ireland.

3. Scott Sumner movie reviews, noting I think he is underrating Return to Seoul, EO, and Pillow Talk.  Usually I agree with Scott 100%.

4. Did the Counter-Reformation harm science?

5. Everybody is the main character.  Or is “leadership that which is scarce”?

6. Why did human society make so little progress for 300,000 years?

The Arrow Replacement Effect and the Dynamics of US Inventors

Ufuk Akcigit and Nathan Goldschlag (my co-author and former student) have an important new paper on the employment and invention dynamics of US inventors. Amazingly they link data on inventors from patents to census data using anonymized, person-level identifiers, known as Protected Identification Keys (PIKs) so they also have individual data on earnings and employment and they link that data to data on firms.

Ultimately, we observe the employment histories of approximately 760 thousand inventors associated with 3.6 million patents granted between 2000 and 2016.

What they find is twofold. First, an increasing number of inventors are being hired by large incumbent firms (left below). Second, when inventors move to large incumbent firms they earn more but they invent less, compared to similar inventors who go to young firms (right below). Why would an incumbent firm pay more for less productive workers? One possible answer is the Arrow replacement effect, namely a monopolist has less incentive to innovate than a competitive firm becasue the monopolist has a bigger opportunity cost, namely it’s own profits. As Arrow put it: “The preinvention monopoly power acts as a strong disincentive to further innovation.” A logical extension is that a monopolist will be willing to pay not to innovate and one way of doing that is to hire inventors who, if they worked at an entrant, would threaten their monopoly profits.

This is an important paper on declining dynamism in the US economy.

Addendum: In a second paper they use their extensive data to discuss the demographic characteristics of inventors.

Khan Academy Joins with OpenAI

One model of a future course is a super-textbook: lectures, exercises, quizzes, and grading all available on a tablet with artificial intelligence routines guiding students to lectures and
exercises designed to address that student’s deficits and with human intelligence—tutors—on call on an as-needed basis, possibly for extra marginal fees.

That was Tyler and I in our 2014 paper. Here’s the Washington Post on the Khan Academy and OpenAI colloboration.

…last week, the private Khan Lab School campuses in Palo Alto and Mountain View welcomed a special version of the [GPT] technology into its classrooms.

Rather than solve a math problem for a student, as ChatGPT might do if asked, Khanmigo is programmed to act like “a thoughtful tutor that’s actually going to move you forward in your work,” says Salman Khan, the technologist-turned-educator who founded Khan Academy and Khan Lab School.

Khanmigo was developed in concert with OpenAI, the nonprofit tech start-up that created GPT-4, the underlying technology for the latest version of ChatGPT. OpenAI did not respond to a request for comment on the partnership.

How to visit Italy

Ajit requests such a post, and I note that plenty of people have plenty of experience with this topic.  So I’ll offer a few observations at the margin:

1. Venice, Florence, and Rome have, on average, the worst food in Italy.  They have some wonderful places, but possibly hard to get into, requiring advance planning, and often expensive.  For random meals, those cities are not impressive, noting that Rome, due to its size, is much better than Venice or Florence.

2. My favorite “single sights” in Italy, moving beyond the core sights of Rome, Venice and Florence, are the Giotto chapel in Padua, the Basilica in Ravenna, and the Cathedral in Monreale in Sicily (near Palermo).  To this day, they remain underrated sights.  As for the major cities, both Genoa and Torino are underrated.

3. My favorite food in Italy would be in Sicily, Naples, and the lower-tier towns of the North, such as Bologna and Parma.  The area near Torino/Piedmont would be another contender.  I have heard Veneto is wonderful for food, though have only had a single meal there, which was indeed outstanding.  In Sicily, don’t order the usual Italian dishes (which are available and excellent), rather look for regional offerings which reflect the area’s Arabic heritage.  Orange slice and mint — bring it on!

4. Usually there is little gain from pursuing Michelin-starred restaurants in Italy.  You want the “two-forker” places with outstanding regional cuisine.  Originality, which is rewarded by the Michelin system, too often is a negative in Italian food.

5. Italy has a large number of third-tier towns which are wonderful for walking through.  But you don’t need to overnight in them, so there is much to be said for randomly driving around Italy, but avoiding the larger cities.  Stop, walk for a few hours, take a meal, and then move on.

6. There is a great deal of available trip prep material for Italy in the form of movies, fiction, and history.  Most of all, however, you should focus on using picture books to have an advance sense of the art and architecture.  The classic book on Italy, Luigi Barzini’s The Italians, is still worth reading.  And often the postwar fiction, or even Manzoni, are better trip prep than the very famous classics such as Dante and Petrarch (though you should read them anyway, but for other reasons).

What else?

The game theory of an AI pause

The issues go well beyond China:

What might countries such as Israel or Japan do if their most important ally decides to pause work on AI? Might this not lead to a proliferation of GPT-like models across more countries — exactly what the pause advocates were trying to avoid?

And if the goal is to “Pause Giant AI Experiments,” which is what the letter is titled, what of smaller ones? What if a small company has an ongoing experiment but is nowhere close to having an effective product? A six-month suspension would damage their future business prospects and serve to entrench the incumbents. What if one of those new upstarts is working to come up with very good safety and alignment procedures? What if its AI might help cure cancer?

There is little evidence that proponents of a delay have thought through the major secondary effects of their 600-word proposal. Maybe they could have made a stronger argument if they they’d had more time to prepare — say, another six months?

Here is my full Bloomberg column.

A new paper on infrastructure costs

Despite infrastructure’s importance to the US economy, evidence on its cost trajectory over time is sparse. We document real spending per new mile over the history of the Interstate Highway System. We find that spending per mile increased more than threefold from the 1960s to the 1980s. This increase persists even conditional on pre-existing observable geographic cost determinants. We then provide suggestive evidence on why. Input prices explain little of the increase. Statistically, changes in income and housing prices explain about half of the increase. We find suggestive evidence that the rise of “citizen voice” in government decision-making increased spending per mile.

That is from American Economic Journal: Applied Economics, by Leah Brooks and Zachary Liscow.

Does natural selection favor AIs over humans? Model this!

Dan Hendrycks argues it probably favors the AIs, paper here.  He is a serious person, well known in the area, home page here, and he gives a probability of doom above 80%.

I genuinely do not understand why he sees so much force in his own paper.  I am hardly “Mr. Journal of Economic Theory,” and I have plenty of papers that you could describe as a string of verbal arguments, but here is an instance where I would find an actual model very useful.  Evolutionary biology is full of them, as is economics.  Why not apply them to the AI Darwinian process?  Why leap to such extreme conclusions in the meantime?

Here are two very simple ideas I would like to see incorporated into any model:

1. At least in the early days of AIs, humans will reproduce and recommend those AIs that please them.  Really!  We already see this with people preferring GPT-4 to GPT 3.5, the popularity of Midjourney 5, and so on.  So, at least for a while, AIs will evolve to please us.  What that means over time is perhaps unclear (maybe some of us opt for ruthless?  But do we all seek to hire ruthless employees and RAs?  I for one do not), but surely it should be incorporated into the basic model.  How much ruthlessness do we seek to inject into the agents who do our bidding?  It depends on context, and so is it the finance bots who will end the world?  Or perhaps the system will be tolerably decentralized and cooperative to a fair degree.  If you are skeptical there, OK, but isn’t that the main question you need to address?  And please do leave in the comments references to models that deploy these two assumptions.  (With the world at stake, surely you can do better than those bikers did!)

2. Humans can apply principal-agent contracts to the AI (again, at least for some while into the evolutionary process).  Keep in mind if the AIs are risk-neutral (are they?), perhaps humans can achieve a first-best result from the AIs, just as they can with other humans.  If the AIs are risk-averse, in the final equilibrium they will shirk too much, but they still do a fair amount of work under many parameter values.  If they shirk altogether, we might stop investing in them, bringing us back to the evolutionary point.

Neither of those points are the proverbial “rocket science,” rather they are super-basic.  Yet neither plays much if any role in the Hendrycks paper.  There are some mentions of various points on for instance p.17, but I don’t see a clear presentation of modeling the human choices in a decentralized process.  p.21 does consider the decentralized incentives point a bit more, but it consists mostly of two quite anomalous examples, such as a dog pushing kids into the Seine to later save them (how often?), and “the India cobra story,” which is likely outright false.  It doesn’t offer sound anecdotal empirics, or much theoretical analysis of which kinds of assistants we will choose to invest in, again set within a decentralized process.

Dan Hendryks, why are you so pessimistic?  Have you built such models, fleshing out these two assumptions, and simply not shown them to us?  Please show!

If the very future of the world is at stake, why not build such models?  Surely they might help us find some “outs,” but of course the initial problem has to be properly specified.

And more generally, what is your risk communication strategy here?  How secure, robust, and validated does your model have to be before you, a well-known figure in the field and Director at the Center for AI Safety, would feel justified in publicly announcing the > 80% figure?  Which model of risk communication practices (as say validated by risk communication professionals) are you following, if I may ask?

In the meantime, may I talk you down to 79% chance of doom?