Toward New Standards for the Mathematical Community


We are currently experiencing a scientific revolution, which is disrupting our practices and requiring the mathematical community to redefine how it operates. Below, I share a few thoughts on a possible future. I do not claim to be right, nor that these thoughts are original. The main intention is to foster reflection on the redefinition of our community.

General remarks:
1- AI is here, and a fraction of mathematicians will inevitably use it. Whether we like it or not, our community must adapt. The question is not whether researcher A or B will use it, or approve it. The question is how we organize ourselves in a world where AI can, at the very least, produce a PhD thesis by current standards in a matter of weeks.
2- In science, what matters is the collective understanding we gain of a subject. It does not matter whether a result was first proven by A, B, I or an AI. Moreover, a second proof is often much more profound and interesting than the first, a contribution that is currently undervalued. What counts is how we digest a subject and gain a deep understanding of it (in particular, books are important and frequently under-appreciated).
3- The question of individual achievements only arises because we need to grant funding, and specifically jobs, to people. Today, hiring relies heavily on the “production of new results.” In my opinion, this criterion is used for several reasons: (1) it has a certain objectivity (easier to measure than deep understanding of a subject), (2) it scales (one can quickly pre-select among hundreds of applications), (3) until now, it was quite well correlated with a deep understanding of a subject. Obtaining new and difficult results could serve as a proxy to assess that depth of understanding. Today, this correlation is broken: a hard problem can be solved with the help of AI by someone who has no understanding of the field.
4- I therefore believe we can collectively adapt to these changes if we redefine our goals and make sure the new generation continues to develop its cognition (by thinking through problems on their own, learning from errors, etc.). Immediate gains (proving lots of low-hanging fruit results) must not come at the expense of the long term. For this, the production of a “first proof” must no longer be an evaluation criterion; we must return to the fundamental task of evaluating mathematical depth.
5-  A primary focus must be on our expectations for a PhD thesis. A PhD plays a triple role: developing the student’s cognitive abilities, assessing their research potential, and discovering / contributing to a field of study. Until now, we managed to combine these three aspects by writing suitable research papers, typically in the style of what AI can currently produce. With unlimited access to AI, this synchronization seems to be getting lost, as we cannot expect a young student to have the maturity needed to "surpass" an AI.

Ideas / proposals:
1- To evaluate the depth of a young researcher, one needs to talk with them for 2 to 3 days. This is doable for 10 people, but it does not scale (until now, publications served as an initial filter). I think this is the core of the problem and the most pressing issue to solve. One option is to ask their PhD or postdoc advisor for their opinion, but we would quickly fall back into nepotism and cronyism. Another option is to adopt a very different organization. For example, something like:
  - People working on a subject would work in an open and collaborative way. The focus would not be on obtaining new results per se, but on increasing the understanding of a subject. A newcomer would be judged by this small community, which would appreciate their contribution. The difficulty lies in ensuring the field's leading figure does not hold too much influence...
  - Each department would then assess the advisability of developing a specific direction and ask for feedback from the community working on that subject. The difficulty is preventing departments from becoming thematically insular, by remaining focused solely on their own communities.
2- Journals must redefine their goals. Since automatic verification of papers can now be done almost automatically, simple repositories for « first proof » papers, acting as databases, can be sufficient as long as (1) there is a medium for sharing the highlights with the community, and (2) search engines can find them. The main role of articles (which I imagine as a collective output of a community) could be to communicate advances in the understanding of a field, highlighting salient and deep points, putting phenomena into perspective, explaining some (truely) innovative proof techniques, and relegating non-fundamental technical aspects to some « compiled » appendices, or external supplementary materials. The articles could also propose some « program » for solving some specific questions. They should not serve as the recognition of a personal achievement.
3- The nature and assessment of PhD thesis must evolve. An initial, well-framed thesis project makes it possible to learn how to organize and prioritize one’s thoughts, build a long and complex strategy of proof, engage in critical assessment of our own work, write a paper, and understand what it entails. It is also an opportunity to experience one's own ability to gradually construct a long proof (so daunting at first) while simultaneously observing one's own limitations. One idea (among many others) could be to have several phases during the PhD program. An "initial" phase, typically lasting one year, dedicated to formulating and writing an unambitious result without AI, which would not be intended to appear in the final dissertation. This would be a pure academic exercise, with the thesis advisors acting as guarantors of the process through active supervision. Following this phase, a "research" phase would begin with the possible assistance of AI, but with a degree of "regulation" regarding its use in order to foster a deep understanding rather than a superficial grasp of the subject.
4- We cannot let a handful of private companies monopolize AI for mathematics. We need to develop a collective and sustainable instrument serving the academic community (and possibly broader audience, along the lines of scikit-learn). The investment is substantial, but physicists know how to do this well. Current models are extremely large, but we can build smaller models, highly specialized in mathematics, that do not provide chocolate cake recipes, advice on how to organize one’s day, or replies to emails. There is no need to memorize all that information, which allows us to reduce model size. However, the reasoning must be top-tier. A first step (which appears achievable) is to launch a European project dedicated to developing the « reasoning layer » on top of existing open-weight models. It would enable us to gain expertise, and ultimately develop open frontier AI for mathematics in the long run. We must manage to unite at the European level around such a project (like Airbus, CERN, ITER, etc.).