postMy 2 cents on LLMs and scientific publication

At ACT2026 we had an interesting community session to reflect on the impact of LLMs, which set in motion some reflections which I would like to voice out loud. I would love to hear other people's opinion too.

When it comes to work submitted to a conference or journal, my position is that we should defer to the author's accountability for what they write, and we should first and foremost judge their submission by the quality of its content.

There are exceptions to this rule, clearly we would not want work produced in harmful way to be allowed (e.g. by forcing a student to write, or plagiarizing someone's work, using unethically sourced data, etc.). Some people might consider any LLM usage to be complicit in the unethical practices of companies providing the service, and thus the matter of what we consider `harmful production' of scientific work is a political one.

The question to focus on is then: what does an effective LLM policy look like? Bans or forced disclosure rules cannot be enforced and thus rely on trusting those who are subjected to them. But if we trusted authors to begin with, then we would not need the rules in the first place---unless we expect that such bans would effectively force otherwise honest people to stop using LLMs altogether, which seems unlikely. What is much more probable is that people will be dissuaded from disclosing usage altogether.

On this topic, I cannot recommend enough this video essay by Dr. Fatima where she discusses whether shaming or outright exclusion is an effective strategy to change people's behaviour (spoiler: no). This has been studied a lot and she interviews a researcher in the field. Rather, education, `harm reduction', and positive reinforcement of ethical practices works much better in effecting behavioural changes amongst peers.

Reviewing is a very different matter though, since the goal there is not to produce a high-quality artifact but to engage in a high-quality process. Reviews are a communal process of digestion of the scientific production, and reviewers are chosen for their taste and know-how, shaped by continuous and personal engagement. Going to conferences, seminars, reading and writing, conversating with other members of the community, all these activities are the field itself, and thus every issue of relevance cannot be addressed without being immersed in such an environment. This is simply something LLMs are not doing, especially the publicly available ones. Asking Claude to review a paper would be like asking an exceptionally well-read random person off the street, and that would be a totally silly way of judging scientific work.

To be clear, I believe there are many tasks that are fine to delegate to the LLM: find typos, find missing related work, find obvious flaws in the arguments of the paper, etc. This is using LLMs as tools and involve no delegation of judgment.

Since reviewers are chosen among the active members of the community, I find it unlikely that enforcing a no-LLM reviews policy would be as hard and pointless as for submissions. I like to review papers! I am delighted to be asked my professional opinion about a work, and I expect this sentiment to be shared. Writing reviews does not result in academic beans, so there is little incentive for people to go out of their way to violate a rule. Thus, I believe policing reviewers to be a much less adversarial process.

The biggest LLM-related issue facing conferences at the moment concerns the asymmetry between reviewing and submitting. As LLMs get better at mimicking human reasoning, it becomes incredibly hard to judge the correctness of a paper. Note this has little to do with LLMs' actual capabilities: everyone should agree that what LLMs write has all the characteristics of human prose but does not stand similar thinking processes, and thus has very different failure modes that are not easily detected by our untrained brains.

A toy example comes from the usage that mathematicians make of words like `easy', `trivial', `straightforward', `routine', etc. We are trained to use those for proofs we checked but would be too uninteresting to report (some people think this is a bad habit, but I disagree). We are also accustomed to a roughly similar attitude in other mathematicians, so one can `trust' such a judgment. But if an LLM writes it, who do we trust? Aside from having a completely different thinking process, the LLM has not been socialized as a mathematician and has no reputation to defend. So the same signifiers that we learned to interpret coming humans suddenly take on a close but different meaning when used by an LLM.

This is only half of the problem, though, since, as I said above, one should judge the paper by what it's written on the page, and review it by filtering the author's prose through one's own brain and reproducing the thought process. In this sense, it does not matter if the LLM thinks differently than us, as long as its output can evoke correct reasoning in our minds---that's really what we check when we review mathematical prose anyway.

The other, more worrying half of the problem is that LLM-generated papers are poised to become very cheap to generate while reviewers' time is not getting any easier to come by. How to cope with this challenge?

AI is exposing and straining those systems that treat humans like machines. We should react by stopping that and instead turn up the human component to eleven. Thus rather than treating the review process as a rubber stamping exercise, let's open it, make it more deliberate, participated, and ultimately a faithful reflection of the scientific conversation and community it should reflect.

The single biggest step in this direction would be adopting open reviews---this is not even new, many AI conferences do this already. Papers are assigned to primary reviewers whose job is to start the conversation by ensuring that at least some reviewers thoroughly assess the work. Their analysis should then be open to other reviewers and authors, all of whom should be able to vote and respond to others' comments (a flexible pseudonymity system could guarantee accountability without sacrificing anonymity when people want it). The review process becomes a conversation, a living and transparent record of the scientific discussion pertaining a work.

You can comment on this post on Mastodon.