Home » General » AI-Assisted Tools in the CHI 2027 Papers Review Process: What They Do, What They Do Not Do, and How Humans Remain Responsible

AI-Assisted Tools in the CHI 2027 Papers Review Process: What They Do, What They Do Not Do, and How Humans Remain Responsible

AI-Assisted Tools in the CHI 2027 Papers Review Process: What They Do, What They Do Not Do, and How Humans Remain Responsible

Authors: Anna Cox, Tony Tang, Thomas Kosch, Erin Solovey, Petra Isenberg, Regan Mandryk
Date: 2026-08-29

Table of Contents

Introduction

In our recent post explaining the CHI 2027 Papers review process, we said we would pilot several AI-assisted and automated tools. We also said we would provide more information about those tools. This blog post aims to serve that purpose by covering questions such as: What exactly are these tools doing? What information do they receive? How have they been tested? What happens when they are wrong? Could an AI system reject a paper? Could it disadvantage qualitative, critical, interpretivist, design-led, or otherwise less conventionally structured work? What happens to the content of an unpublished paper when it is processed? What does “piloting” mean in a review process where the consequences for authors are real?

The most important point is this:

No AI system decides whether a CHI 2027 paper is accepted or rejected. Humans retain responsibility and authority for every consequential decision about a submission.

The tools described below provide information, surface possible issues, or suggest possible matches. They do not replace the scholarly judgment of Associate Chairs (ACs), Subcommunity Chairs (SCs), reviewers, or Papers Chairs.

We recognise that simply saying “human in the loop” is not sufficient. Human involvement only provides a meaningful safeguard if we are clear about what the automated system produces, who sees it, what they are expected to do with it, and what happens when human and automated assessments disagree.

This post therefore, describes each tool using the same questions:

  • What problem are we trying to solve?
  • What does the tool do?
  • What does it not do?
  • Where does human judgment enter?
  • How have we tested it?
  • What are its known limitations and safeguards?

We also address broader questions about data, consent, bias, accountability, and the process by which these tools were introduced.

Why use automation at all?

Before describing the individual tools, we want to make one broader point.

The relevant comparison is not between an imperfect AI-assisted process and a perfect human process.

The existing human review process is not perfect.

CHI depends on enormous amounts of expert volunteer labour, and much of that labour is excellent. But people miss things. Reviewers sometimes have incomplete expertise. Papers can be poorly matched. Reviews vary in quality. Different parts of the programme can develop different standards. Administrative problems can go unnoticed. And, at the scale at which CHI now operates, asking people to manually inspect everything creates its own risks.

The historical evidence already shows some of those pressures. CHI 2026 received 6,730 complete submissions, reflecting very rapid growth in submissions. The review restructuring is intended in part to preserve careful scholarly judgment by making more effective use of the finite expert attention available to the conference.

We have also seen concrete examples in which automated checking found problems that had survived human peer review.

For example, GPTZero has been analysing conference proceedings for hallucinated references and has identified cases of references that appeared not to correspond to real publications in venues like NeurIPS and ICLR. These were references from papers that had already passed through human peer review.

This does not demonstrate that an automated system is a better peer reviewer. It demonstrates something more modest, but important: humans and automated systems fail in different ways. A tool that reliably performs a narrow check may help a human notice something that they would otherwise miss.

We have seen the same phenomenon during testing of the desk-reject support tool (more on the testing below).

Our aim is therefore not to automate scholarly judgment. It is to ask whether carefully bounded automated tools can help people perform particular parts of a very large review process more consistently and with better information.

That approach is consistent with CHI’s commitments to transparency, inclusion, quality, and long-term adaptation of the conference. It also means that we have an obligation to examine the limitations of these tools rather than treating their outputs as authoritative.


Tool 1: Automated submission completeness checking

We will be deploying this tool in the CHI 2027 review process. 

What problem are we trying to solve?

CHI submissions now require a substantial amount of information in addition to the paper itself. For CHI 2027, this includes, among other things, the information required under the Review Responsibility Policy and descriptors used to help identify appropriate reviewing expertise.

Some of these requirements are mechanical rather than scholarly. It is better for authors to discover a missing or malformed piece of information immediately, while they still have an opportunity to correct it, than for a problem to emerge later in the process.

What does the tool do?

After submission, an automated completeness check verifies that the required information is present.

The process described for CHI 2027 includes checking:

  • whether review-responsibility slots have been properly declared;
  • whether named reviewer-authors have valid ORCIDs and DBLPs (or N/A is included for DBLP if the author doesn’t have one); and
  • whether sufficient descriptors have been supplied to support keyword-based matching.

These checks are rule-based deterministic checks using the structured submission data from PCS, ORCID, and the DBLP dataset from Aug 2026.

Authors receive a completeness report and have a short period to address flagged issues before screening begins.

What does it not do?

The completeness checker does not determine whether a paper makes a valuable research contribution.

It does not decide whether the methods are good, whether the contribution is novel, whether the argument is persuasive, or whether the paper should ultimately be accepted.

No LLM is involved in this check.

Where is the human decision?

The goal is to ensure that submissions are following the Full Paper Review Responsibility Policy. The tool is checking verifiable metadata, where the purpose is principally to alert authors and organisers to missing or inconsistent information.

How have we tested it?

The compliance logic has been unit tested, and we tested the tool with a test instance of the PCS database. Since submissions on PCS opened, we have also been regularly testing with completed submissions into PCS to check whether the tool correctly identifies problematic reviewer-responsibility tags. Based on simulations of keyword-based matching in universes of 10000 submissions, we found that when reviewers select about eight keywords for their expertise descriptors, this resulted in effective keyword matching.

Known limitations and safeguards

Identity resolution relies on two self-reported identifiers, ORCID and DBLP. We ask for both, because some scholarly records exist in one but not the other. We have identified two failure modes for this:

An author could supply an ORCID or DBLP link that resolves but is not actually theirs. We have no independent way to verify a supplied identifier against the person submitting it; we treat a resolvable identifier as trustworthy, on the assumption that fabricating one is both unlikely and against an author’s own interest.

An author could declare no DBLP profile when one exists. We still attempt to locate a profile independently in that case, but if that search also fails, we assume the author does not have a DBLP profile and is acting in good faith.

Low-confidence or unresolved identity never triggers an automatic consequence beyond notifying the authors while there is still an opportunity to correct them. There will be a small window of opportunity post-submission date to resolve these situations.

Tool 2: Matching papers to ACs and reviewers

We will be deploying this tool in the CHI 2027 review process. 

What problem are we trying to solve?

Good peer review depends heavily on good matching.

CHI is both very large and unusually methodologically diverse. A paper may need expertise in a particular domain, a particular method, a particular theoretical tradition, or some combination of these.

At CHI’s current scale, manually inspecting thousands of papers and potential reviewers is increasingly difficult. The CHI 2027 process, therefore, uses matching information to help make that search tractable.

For several years, ACs have been able to use functionality built into PCS to identify potential reviewers. The algorithm PCS uses is the Toronto Paper Matching System, but ACs found the recommendations unreliable. Our approach to this problem is to try to develop a more comprehensive keyword hierarchy that authors can use to tag their papers, and for reviewers to clearly articulate expertise.

What does the tool do?

Authors provide descriptors that indicate the expertise they believe is required to review their paper. Reviewer and AC profiles provide information about their expertise.

The matching system uses this information to identify potentially suitable ACs and reviewers from the pool of people submitted in the volunteer slots while also taking into account conflicts and workload.

For external reviewing, the process currently described is explicitly advisory: the system suggests candidate reviewers to the AC. The AC chooses the reviewers. The AC can accept a suggestion, reject it, replace it, or identify a different reviewer whose expertise is a better fit (who may or may not be in the pool of people submitted in the volunteer slots).

What does it not do?

A high matching score does not mean that a person is automatically the right reviewer.

A low matching score does not mean that someone lacks relevant expertise.

The system does not know a paper’s intellectual context in the way that an expert member of the community might. It is a search and recommendation aid.

Where is the human decision?

The AC remains responsible for assembling a reviewer team that collectively has the topical and methodological expertise needed to assess the work.

This is important because CHI’s minimum-qualification policy explicitly treats reviewer-team expertise as a collective property. ACs are expected to consider whether the team covers both topic and methodological expertise rather than simply selecting three individually plausible names.

How have we tested it?

Matching has been evaluated on simulated worlds: generated papers’ keyword selections, and reviewer profiles with a known “correct” match built in. We stress tested this in various ways (e.g. an author uses very few descriptors or a reviewer profile uses very few descriptors), and observed that the algorithm behaved in sensible ways. We have not yet assessed this against historical CHI assignment data or expert human judgment of match quality.

The matching uses IDF weighting (similar to how classic information-retrieval systems behave), so rarer, more specific keywords count for more in a match than common terms. A paper tagged with an uncommon method or subfield is matched more precisely toward the smaller pool of people who share that specific expertise, rather than being pulled toward generalists who happen to share the paper’s more common tags.

Known limitations and safeguards

Matching systems can reproduce biases present in expertise descriptions, publication histories, keywords, and other input data.

They can also miss unusual combinations of expertise, which are particularly important at CHI.

For that reason, a match is a recommendation to a human AC, not an assignment that the AC is required to accept.

Based on the actual assignments that ACs make in CHI 2027, we can assess whether the keyword matching for recommending reviewers was useful to ACs. We will provide this as advice to the CHI Steering Committee moving forward.

Tool 3: Desk-reject support

We will be deploying this tool in the CHI 2027 review process. 

What problem are we trying to solve?

Traditional desk rejection deals with submissions that cannot proceed through peer review because they fail basic conference or ACM requirements.

Examples described in the CHI 2027 process include incomplete submissions, anonymisation violations, incorrect format, undeclared concurrent submissions, papers outside CHI’s scope, submissions that cannot meaningfully be reviewed, and violations of relevant ACM policies.

At thousands of submissions, finding all such cases manually is itself substantial work. For CHI 2026, ACs were asked to manually check each reference to ensure reference integrity—this is a task easily automated. With more than 600 ACs, ensuring that these criteria are applied consistently has also been difficult.  

What does the tool do?

The screening tool surfaces candidate submissions for inspection by the programme committee.

That wording is deliberate.

It is effectively saying to a human: there may be something here that you should look at.

What does it not do?

It does not desk-reject the paper.

A flag is not a decision.

The tool is not authorised to turn its assessment into a rejection outcome.

Where is the human decision?

The CHI 2027 process contains multiple levels of human verification.

A flagged paper is examined by an AC. If the AC believes desk rejection is appropriate, this must then be confirmed by an SC before the decision is issued. If human decision-makers disagree with the automated flag, the flag does not determine the outcome.

This is an example of where we intentionally use automation asymmetrically: the tool can direct human attention; it cannot exercise the authority attached to that attention.

How have we tested it?

The checks were developed and measured retrospectively against the CHI 2026 cycle. Deterministic checks were measured on the full analysis corpus of 6,286 submissions (6730 minus conflicted submissions for the tool developers). Checks that involve a model were measured on an enriched sample of 744 papers, comprising all 394 desk-rejected submissions, 250 accepted papers and 100 rejected papers, so that rare categories appeared in sufficient numbers to be measurable.

No rule check fired on any of the 110 CHI 2026 award-winning papers included.

Individual checks performed as follows: 

  • Anonymisation. From the text alone, the check identifies 26/71 real anonymisation desk-rejects (37%), with no false positives among 250 accepted papers. The missed cases are largely structural rather than a tuning problem: according to the rationales, 31% of real breaches leak through self-citation, and others through figures and supplementary files that the tool cannot see. 
  • Template. 19 submissions were flagged for using a two-column rather than single-column review format. All 19 had in fact been desk-rejected by CHI. 
  • Wrong document type. Three flags, all three on desk-rejected papers.
  • Duplicate submissions. Every match found across the corpus was a genuine duplicate confirmed by the chairs.
  • Masked references. Fires on 1.0% of the held-out test split (15/1,555), all model-confirmed: 12 genuine cases and 3 correctly cleared.
  • Scope. On the training split, the flag tier reached 100% desk-reject precision, touching none of 601 accepted papers; on the held-out test split it again touched none of 392 accepted papers. Recall, however, was 21% on test against 55% on train. The papers it misses are ones the chairs scoped out on audience-fit grounds despite a literal human-interaction component, which the tool deliberately does not attempt to predict.

This last result is worth highlighting. The check is a detector of one unambiguous subtype of scope problem, and nothing more. Reporting it as a general scope detector would overstate it considerably.

Known limitations and safeguards

Some desk-reject criteria are straightforward to detect; others require interpretation.

The checks the tool attempts are: author and institution anonymisation; anonymisation of supplementary and external links; references masked as “Anonymous”; use of the single-column review template; wrong document type, such as an uploaded thesis or a journal-formatted manuscript; reference integrity; whether the submission is reviewable as English text; compilation defects; duplicate submission within the cycle; hidden text aimed at AI readers; and a narrowly drawn scope check.

Several of these are treated as discretionary and never halt a submission on their own. Masked references are an example: four accepted papers were found to carry genuine masked references and were accepted anyway. Compilation defects are likewise advisory, because seven accepted papers carried them.

The tool does not attempt to verify ethics or IRB compliance. Confirming an IRB protocol identifier would require exactly the identifying information that anonymisation policy asks authors to remove, and a check that cannot be performed without undermining the policy it serves should not exist. It also does not attempt undeclared concurrent submission to venues outside CHI, and it does not assess whether a paper is scientifically sound.

The safeguard for ambiguous cases is human examination and multi-level confirmation.

Example Tool 3 Output

Figure 1 shows an example of the tool’s output. The evidence & reasoning provides additional details on why the rule was flagged and where ACs should look in the paper to make their judgement. In this example, the paper violated the anonymization policy, and there were two references that needed manual checking by an AC, as they could not be resolved.

Screenshot of Tool 3 output, "Rule checks — potential policy violations". Introductory text explains that each card links to the CHI/ACM policy it is grounded in, and that a flag marks a potential violation for a human to verify rather than a predicted decision.

The first card, marked with a red cross and labelled RV-3 · MASKED-REFS "Anonymous references", is tagged "Blocking" and "FLAG". It reports that one or more references appear to use "anonymous" or "removed for review", which the anonymisation policy treats as grounds for desk rejection. A fast model reviewed the single candidate and confirmed it as a masked reference. Footer: "deterministic + model-confirmed".

The second card, marked with a question mark and labelled RV-6 · FAKE-REFS "Reference integrity", is tagged "UNVERIFIED". It reports 56 of 56 references checked: 54 resolved, one weak but found, none missing, one that could not be checked, no errors, and one anonymised reference set aside under RV-3 — an incomplete result rather than a clean one, easily verified by hand. Footer: "resolver".

Both cards have collapsed disclosure triangles reading "what this check does & why" and "evidence & reasoning".

A final card with a green tick reads: "8 checks clear — RV-1 · IDENTITY, RV-2 · LINKS, RV-4 · TEMPLATE, RV-5 · NOT-A-PAPER, RV-7 · LANGUAGE, RV-8 · BUILD-DEFECTS, RV-10 · PROMPT-INJECTION, RV-11 · SCOPE".

Figure 1. Screenshot of a rule-check report panel titled “Rule checks — potential policy violations”, showing two flagged cards (anonymous references, reference integrity) and a green line noting eight further checks passed.

Tool 4: AI-assisted rubric report for Assisted Desk Reject (not used in CHI 2027 peer review)

Of all the tools described here, this is likely the one that will generate the most interest.

Unlike checking whether required information has been supplied or suggesting possible reviewers, this tool examines aspects of the content of a submitted paper. In developing this tool we have considered important questions like whether an AI system can assess scholarly work reliably, whether particular research traditions could be disadvantaged, whether an automated report might anchor an Associate Chair’s judgment, and how much influence the tool has over whether a paper receives full peer review.

We want to be particularly clear about the last point:

The ADR tool does not decide whether a paper is desk rejected. It cannot reject a paper. The Associate Chair is responsible for reading the paper and deciding whether to recommend an ADR outcome, and that recommendation is subject to further human oversight.

The purpose of the tool is to give the AC additional, structured information to consider while reading the paper. Its output is advisory and is not being used by ACs for CHI 2027, but is being tested. 

What problem are we trying to solve?

Every paper sent to full review requires three external reviews as well as substantial work from an AC managing those reviews, reading the paper, facilitating discussion, and producing a meta-review. With CHI’s submission volume continuing to grow, sending every submission through that process creates an increasing burden on the people whose expertise makes peer review possible. 

Assisted Desk Reject (ADR) is the language used by ACM to describe rejections based on the judgment of the editor or subeditor (SC and AC in our case) that a paper is either out of scope or so far from acceptable as to make external reviews unnecessary. The ‘Assisted’ in Assisted Desk Reject refers to the assistance that a subchair (AC or SC) is providing to the paper chairs to desk reject a paper, not to any assistance from an AI tool. 

The purpose of the ADR phase is therefore narrower than deciding whether a paper is ultimately “good enough for CHI”. It is to determine whether a submission has a sufficiently realistic path towards acceptance that committing the resources required for full external peer review is appropriate. That distinction is important. The ADR process’ stated purpose is not to produce an overall quality judgment, but to ask whether the paper has sufficient quality to move forward into full peer review. There is a short window of time for the ADR process before review requests go out. With the projected number of submissions for CHI 2027 and beyond, it will be a challenge for ACs and SCs to identify these cases in the time available.

This is also why the tool should not be thought of as an automated paper reviewer. It is intended to support a first-stage assessment. Papers that proceed receive the much deeper evaluation provided by qualified external reviewers and the AC.

What can this tool do?

The rubric was deliberately designed to work across CHI’s methodological and disciplinary diversity. Rather than specifying that every paper must, for example, have a particular sample size, use a particular form of validation, report a particular quantitative metric, or follow a single model of research practice, it focuses on whether the claims being made are appropriately supported by the evidence, argument, analysis, design work, or other form of scholarly contribution being presented.

The AI-assisted implementation makes that broad rubric more concrete through a series of narrower checks. These include questions about:

  • whether the paper provides sufficient grounding, methodological information, evidence and data for its claims to be meaningfully assessed;
  • whether important claims are supported by the evidence or argument presented;
  • whether the core contribution can be understood from the paper itself rather than depending excessively on material elsewhere;
  • what type or types of contribution the paper appears to make and whether the form of validation is appropriate to those contribution types;
  • whether the paper engages with relevant HCI literature;
  • aspects of reference integrity and reference quality;
  • whether the paper’s structure and other characteristics are unusual relative to comparable CHI papers; and
  • whether information that would normally be needed to assess the claimed contribution is actually stated.

Some of these checks are deterministic: they involve measuring or looking up something that can be examined without asking a language model to make a scholarly judgment. Others involve model judgment. The tool documentation explicitly distinguishes between the two rather than treating everything as a single AI assessment.

That distinction matters. Counting references, checking for particular structural features, or resolving a citation against bibliographic databases is a very different task from assessing whether the evidence presented adequately supports a scholarly claim.

The tool does not simply ask an AI whether a paper is “good”

One concern people might have is that the system might effectively submit the PDF to a language model and ask, “Is this a good CHI paper?” 

That is not how the tool has been designed.

In particular, we have deliberately not attempted to automate every one of the five ACM criteria in the same way.

Although the ADR framework is organised around ACM’s five criteria, the implementation currently contains no direct Originality score. Novelty is also treated cautiously: the tool can examine whether a paper engages with the literature against which a novelty claim would have to be understood, but it does not attempt to turn “how novel is this idea?” into a reliable automated score.

This is deliberate.

During development, attempts to measure these kinds of judgments performed poorly. Predicting reviewer disagreement on these dimensions from the paper alone performed at chance, and an embedding-based attempt to measure novelty did not perform reliably. The conclusion was not to automate the unreliable judgment more aggressively; it was not to use that measure.

We think that is an important principle: when we do not have evidence that a particular assistive judgment is useful, the answer should not only be to not rely on it, but also to not present it in the report. We have amended the tool as we have learned how it behaves.

Accounting for different kinds of CHI contribution

Authors might also be concerned that an AI system might implicitly assume that every strong CHI paper looks like a conventional quantitative empirical study.

That would be unacceptable for a conference with CHI’s methodological and epistemological diversity.

The tool has therefore been designed to reason about the type of contribution a paper claims to make before applying expectations about how it should be validated.

For example, an artifact or system contribution should not automatically be criticised for lacking the form of controlled experiment that might be expected for a particular empirical claim. A qualitative contribution should not be assessed against the evidentiary conventions of an unrelated quantitative contribution. A paper may also make several different contributions, in which case the system is designed to consider each against the form of validation appropriate to that contribution.

The inferred contribution type is deliberately expressed tentatively. The report makes the premise visible—for example, that the work appears to be making a particular kind of contribution—so that an AC can immediately see if that premise is wrong. The system is instructed not to turn an uncertain categorisation into an authoritative statement about the paper.

Testing also caused us to change how this component works. An earlier approach effectively raised a concern whenever any one of several claimed contributions appeared weak. Because CHI papers commonly make multiple contributions, that rule fired on 87% of the papers in one evaluation set and did not meaningfully distinguish papers receiving stronger or weaker reviewer recommendations. The rule was therefore changed substantially. The current approach uses a much higher bar for producing a finding, and the documentation explicitly states that the contribution-type component should not be presented as evidence that a paper will be rejected.

What can this tool not do?

There are several things the Assisted Desk Reject (ADR) tool is deliberately not authorised to do.

It does not issue an ADR decision.

It does not automatically reject a submission.

It does not determine the final outcome of a paper that proceeds to external review.

Its output is not a vote.

It does not forecast acceptance.

It is not intended to provide an overall ranking of CHI submissions.

The design guidelines state explicitly that subjective outputs are “Advisory, always.” They are intended to identify something that a careful reviewer or AC might reasonably want to examine, provide the reasoning behind that concern, and report confidence rather than present subjective judgments as established facts. The system is also instructed to favour under-flagging where the alternative would be making an unsupported quality accusation.

That conservative design is intentional. In this setting a false positive matters: incorrectly suggesting that there is a serious problem with a legitimate contribution could influence scarce human attention and, if handled badly, could undermine author trust.

Human judgment is not a final check added after an automated decision. It is the decision-making process.

In CHI peer review processes, the AC is responsible for the paper at the assisted desk reject stage.

The AC reads the paper and the report and forms their own judgment. The report is one source of information available to them; it is not a substitute for reading the submission.

If the AC does not believe that an assisted desk reject is warranted, the automated report does not have the authority to produce one.

If the AC does believe that the paper should receive an ADR outcome, there is further human oversight before authors are notified: the proposed outcome is reviewed by the SC.

The key sequence is therefore:

tool produces evidence and possible concerns → AC reads the paper and report → AC makes their own judgment → SC oversight of the AC’s ADR recommendation.

It is not:

AI evaluates paper → AI score crosses threshold → paper is rejected.

We also recognise that “a human is involved” is not, by itself, a complete answer to concerns about automation bias. People can be anchored by decision-support systems, particularly when they are busy. As a consequence, ACs will be given detailed instructions about independent judgment and automation bias because responsible human oversight depends not only on who formally makes the decision but also on how the information is presented to them.

How have we tested the tool?

We have not treated a plausible-looking AI response as sufficient evidence that a component belongs in the review process.

The development process has used historical CHI papers and outcomes to measure individual checks, with particular attention to false positives as well as the ability to detect historical ADR cases.

Testing has also included accepted and award-winning CHI papers. This is important because a system that appears good at finding historically rejected papers but routinely identifies excellent accepted papers as problematic would not be useful.

One example concerns the tool’s assessment of whether a paper provides enough grounding, methods, data and evidence to be meaningfully reviewed.

An early binary version of that check effectively asked whether there was a concern in these areas. That turned out to be much too permissive: almost every real paper has some limitation that could plausibly be described as a “concern”. In testing, the approach flagged 27% of CHI award-winning papers. That was not an acceptable operating point.

The question was therefore redesigned around a narrower concept: reviewability. Instead of asking whether the paper has any weakness, the tool asks how far it falls below the point at which a reviewer can meaningfully assess its claims.

On the calibration set, the revised measure achieved an Area under the Curve of approximately 0.84 on held-out data. (Area Under the Curve is a plot of true positive rate against false positive rates, the closer the number is to 1, the closer the model is to perfect predictions of a binary outcome.) At the selected operating threshold, it identified about 53% of historical ADR cases, while flagging approximately 6% of award papers. On the held-out split, the check fired on 26.8% of submissions overall.

Those numbers are not presented as evidence that the component is perfect. They illustrate the trade-off we are making deliberately: we would rather leave some possible problems for the AC to identify than generate a large number of false accusations about strong papers.

We have also tested modifications that sounded sensible but did not improve the system.

For example, one experiment provided the reviewability component with explicit contribution-type-specific expectations. The intention was to make the system more sensitive to methodological diversity. In practice, performance became slightly worse and false positives on award papers increased. The likely reason was that the additional information encouraged the model to turn contribution descriptions into a checklist and search for missing elements rather than asking the narrower question of whether the work was actually reviewable.

That version was therefore rejected.

This is an important part of how we are approaching both the iterative development of the tool: changes that sound intuitively better are not automatically retained when the measurements say otherwise.

Other components have been evaluated independently. For example, checks related to unusually low engagement with HCI literature were calibrated against historical out-of-scope desk-reject cases as well as accepted papers, rather than treating any absence of familiar venue names as evidence that a paper is out of scope. The resulting tool raises a question for human consideration rather than declaring that the paper does not belong at CHI.

Similarly, structural checks were evaluated against historical CHI outcomes. Their role is deliberately narrow: to identify unusual characteristics worth a second look, not to assume that every accepted CHI paper should have the same structure. The norms are conditioned on contribution type so that, for example, a design paper is not automatically penalised for lacking a section that would be expected in a conventional empirical paper.

We are also recording where the tool is weak

Testing a tool responsibly means looking for reasons not to trust its output.

The current documentation therefore records known limitations alongside successful measurements.

For example, the reviewability measure described above is partly correlated with paper length. Paper length alone separates some historical ADR papers from accepted papers surprisingly well. Although the model adds some information beyond length, the documentation explicitly warns against describing the resulting score as a pure measure of methodological sufficiency.

Likewise, the contribution-type component has been found useful for explaining what kind of contribution a paper appears to make and what form of validation is relevant, but its verdict does not predict reviewer recommendations beyond what can already be explained by body length. The tool documentation therefore explicitly says not to present that finding as evidence that a paper will be rejected.

Another methodological-detail component was found to identify many genuine absences, but to do so far more exhaustively than human reviewers considered useful. It fired on 85 of 87 papers in one validation analysis. The response was to demote that component to author-facing advice rather than allow it to compete for an AC’s attention as an ADR signal.

We think reporting examples like these is important. A tool used by CHI should not be described only in terms of where it succeeds.

Why compare the tool with the real human review process?

The standard for this pilot should be demanding. But the relevant comparison is not an imperfect automated tool against an imaginary human process in which every problem is detected and every paper is evaluated consistently.

Human review is itself imperfect.

Reviewers sometimes miss important problems. Reviewers sometimes disagree about contribution types or apply inappropriate methodological expectations. People miss reference problems. They can overlook omissions. They can also reach different judgments about exactly the same work.

That is why the ADR tool is being investigated as decision support, rather than as a replacement for reviewers.

We have already seen examples where automated checking can identify narrow problems that humans have missed. For example, the masked references and compilation defects in accepted CHI 2026 papers. 

Those examples do not show that an AI is a better peer reviewer. They show something much more limited and useful: different kinds of checking catch different kinds of errors.

The question we are asking is therefore whether a carefully constrained tool can help an experienced AC notice issues that deserve attention while leaving scholarly judgment with the people responsible for making it.

What safeguards are built into the ADR tool?

Several safeguards follow directly from what we learned during development.

First, the system is advisory. A model-generated concern is not a finding of fact and is not an ADR decision.

Second, we deliberately under-flag. Where subjective interpretation is required, the guidance prioritises avoiding unsupported quality accusations over maximising the number of historical ADR cases detected.

Third, we do not automate judgments simply because they appear in the ACM rubric. Direct originality scoring was omitted, and attempts at automated novelty measurement were not retained when they proved unreliable.

Fourth, contribution type matters. Expectations are not intended to assume a single preferred CHI methodology, and inferred contribution types are presented tentatively rather than as facts.

Fifth, individual components are tested separately. A strong result from one check is not treated as evidence that every AI-assisted judgment works equally well.

Sixth, failed approaches are discarded or demoted. We have examples where testing changed thresholds, changed aggregation rules, removed candidate approaches, or restricted a check to author-facing advice because the evidence did not justify giving it greater weight.

Seventh, the AC remains responsible for reading the paper and making the judgment. The tool supports attention; it does not possess decision-making authority.

Operational Details

There are important operational details that should accompany this description, particularly because they affect confidentiality, participant consent, and the community’s ability to evaluate the pilot.

Models used: The use of the tool is not tied to any vendor: all model access goes through the LiteLLM unified interface (~100 providers), chosen entirely by environment variable. There is no default provider — users pick one explicitly. We tested the model using Claude Open 4.8. 

Processing and retention. Submissions are processed only through zero-data-retention or own-tenancy endpoints. The reference resolver’s web search likewise uses a zero-data-retention backend. In our own testing, caches and analysis corpora were held outside cloud-synced folders.

Adversarial testing. Hidden text aimed at AI readers is a potential threat to the tool’s efficacy. The tool inspects each PDF’s rendering instructions for text a human cannot see — invisible render modes, white-on-white text confirmed by actually rendering the region, sub-2pt type, off-page content — and then checks whether that text contains instructions directed at AI reviewers. Any suspicious hidden text is scrubbed from what every other check reads, and is only ever quoted as untrusted input; it is never followed. Across all the entire CHI 2026 corpus tested, zero prompt-injection attacks were found. About 5.5% of papers carry benign hidden text such as figure text layers and template remnants, and rendering-based arbitration reduces what reaches a reviewer to roughly 0.6%. The detection tier has been tested against synthetic red-team fixtures.

Auditing across research traditions. Expectations are conditioned on contribution type throughout, and structural norms were measured separately by contribution type so that, for example, a design paper is not penalised for lacking a section belonging to a conventional empirical paper. But no separate audit has been conducted across specific research traditions. Contribution-type conditioning is not the same thing as an epistemological audit.

Example Tool 4 Output

Figure 2 shows an example of the output of Tool 4. The evidence & reasoning provides additional details on where ACs should look in the paper to make their judgement. In this example, the paper did not engage with the HCI literature very much and the AC is directed to check the framing to see if it is a contribution to HCI. The tool also suggested looking at whether the contribution warrants the length of the paper and whether authors justified the length.

Screenshot of two sections of Tool 4. 

The first, “Novelty”, explains that the tool does not judge novelty itself but checks its prerequisite: that the submission engages the prior literature any novelty claim would be measured against. Below it sits a card with a lightbulb icon, labelled PF-5 · HCI-LIT “Engagement with the HCI literature”, tagged “Advisory” and “SUGGESTION”. It reports that the paper’s bibliography cites only 2 SIGCHI/HCI references out of 78 (3%), and poses a question for the reviewer or AC: is this work framed as a contribution to HCI? It notes that papers desk-rejected as out of scope look like this, but so do good papers drawing on an adjacent field, so the flag is a prompt to check the framing rather than a judgement that the paper is out of scope. Footer: “deterministic + model-judgment”.

The second section, “Importance”, asks whether the contribution is articulated, present, and proportionate to the length of the paper. Its card, also with a lightbulb icon, is labelled PF-1 · LENGTH “Length vs. contribution” and tagged “Minor” and “SUGGESTION”. It gives a policy-basis word count of roughly 12,465, above CHI’s 12,000-word threshold, and quotes CHI’s rule that submissions above 12,000 words will be desk-rejected if their excessive length is not justified. It advises making the justification clear to reviewers and verifying against the manuscript’s own count, since the tool’s basis is an approximation of CHI’s, and adds parenthetically that the length appears justified by the contribution. Footer: “model-judgment”.

Both cards have collapsed disclosure triangles reading “what this check does & why” and “evidence & reasoning”.

Figure 2. Screenshot of two report sections, “Novelty” and “Importance”, each containing one advisory suggestion card — on engagement with the HCI literature and on paper length versus contribution.

On using the Tool in the CHI 2027 Review Process

The ACM policy on peer review does not prevent reviewers from using off-the-shelf LLMs to support peer review, as long as confidentiality of the submission is maintained. Over the last few years, other conferences have found that a large proportion of their reviews were likely generated by AI. Additionally, other conferences have found that customized and tailored AI systems outperform LLM-generated reviews at detecting weaknesses. In our context, without a customized tool, disciplinary biases will not be addressed, and safeguards cannot be implemented to prevent problematic, inappropriate, or incorrect conclusions. Our tool is designed to account for the range of CHI submission types, to explicitly guard against applying one disciplinary standard to another discipline, and to always present evidence and not conclusions. We believe that our tool is more effective, safe, principled, and tailored to the CHI conference than what reviewers and ACs who use LLMs to support their peer review will receive. 

However, we recognize that a tool used to support peer review should be transparently developed, with time given for community testing and feedback. Our intention has been to release the tool for the community to provide feedback (similarly to how we released the system underlying Tool 2, requesting community engagement and feedback). However, given the very tight timelines (due to the system being developed and tested entirely by volunteers in their spare time and requiring multiple layers of input and approval), we find ourselves very close to the submission date without a public release. What this means is that we will not be using this tool to support ADR processes in 2027. 

A principle running through all tools: assistance is not authority

It is tempting to talk about all of these systems simply as “the AI tool.” That obscures important differences.

Checking whether an ORCID is present is not the same task as suggesting a reviewer.

Suggesting a reviewer is not the same task as surfacing a possible desk-reject case.

And neither is the same as producing a preliminary rubric assessment of scholarly work.

We should therefore evaluate each tool according to the task it actually performs and the consequences its output can have.

But there is one principle common to all of them:

We are not transferring responsibility for CHI’s scholarly decisions from people to AI.

CHI’s review structure assigns responsibility to identifiable people – people who meet the qualifications defined in the Minimum Qualifications For Peer Review Roles Policy – precisely because scholarly judgment requires expertise, interpretation, accountability, and the ability to recognise when a rule, rubric, reviewer, or automated output does not fit the case in front of them.

Human review is a safeguard—but humans also need safeguards

We also want to avoid romanticising the status quo.

CHI has longstanding mechanisms for human judgment, discussion, escalation, and calibration because individual human judgments are themselves imperfect.

The historical review process has shown substantial variation. For CHI 2026, desk-reject and assisted desk-reject rates varied considerably across subcommittees (see figure 2). More broadly, analyses of the Program Committee show differences in experience and reviewing cultures across the conference. When asked at the CHI 26 peer review feedback session whether people had experienced having a paper rejected from a SIGCHI conference and the same paper subsequently accepted (and even having won best paper) at the next deadline, many hands were raised. This common experience demonstrates some of the randomness that is inherent in peer review. The luck of the draw with AC and external reviewer assignment can result in very different outcomes for the same paper. This randomness was also demonstrated in the 2014 and 2021 experiments done at NeurIPS.

Human reviewers sometimes miss factual problems. They sometimes misunderstand methods. They sometimes provide inadequate reviews. They sometimes disagree profoundly.

That is why CHI does not normally ask one person to make every decision alone.

The same philosophy should govern our use of automation.

The question is not:

“Is the AI correct?”

The more useful questions are:

“Does this tool help a responsible AC with HCI expertise notice something useful?”

“What happens when it is wrong?”

“Can people override it?”

“Can we detect systematic errors?”

“Does it improve the overall process compared with the process we actually have—not compared with an imaginary perfect one?”

Those are questions we can test empirically.

Data, confidentiality, and research-participant consent

Several community members have raised an especially important question: what happens when the submitted manuscript contains participant data or quotations and the authors’ ethics approval, IRB documentation, consent forms, or assent forms contain restrictions concerning AI processing?

Authors can never be sure that their published paper is not going to be put through an AI. It almost certainly will. It definitely will if it is published by ACM as papers in the DL are processed by AI to create AI summaries. Reviewers cannot be expected to review a paper that will later be redacted. Therefore, authors should not submit papers that contain information that they are not OK with being published later. None of the tools described provide any training data to any AI companies. They all maintain confidentiality of the submissions. And all tools conform to ACM policies on peer review. 

Which tools read your manuscript. Only the desk-reject support tool reads the submission itself. The completeness checker works from PCS metadata and public bibliographic records, and involves no language model at all.

The model. The desk-reject support tool is a cascade of checks, not a single system. Some are entirely deterministic and involve no model: counting words, measuring page layout, extracting links, matching references against bibliographic indices, detecting compilation defects, and detecting duplicate submissions within the cycle. Others do use a large language model: judging whether identifying text constitutes an anonymisation breach, confirming whether a bibliography entry is genuinely masked, identifying wrong document types, the scope check, adjudicating references that the resolver could not locate, and determining whether hidden text contains instructions directed at AI readers. The distinction is maintained deliberately, and the report shows which is which.

The use of the tool is not tied to any vendor: all model access goes through the LiteLLM unified interface (~100 providers), chosen entirely by environment variable. There is no default provider — users pick one explicitly. We tested the model using Claude Open 4.8. All model-dependent measurements reported in this post are specific to that model and do not transfer to another without recalibration.

Processing and retention. Submissions are processed only through zero-data-retention or own-tenancy endpoints. The reference resolver’s web search likewise uses a zero-data-retention backend. In our own testing caches and analysis corpora were held outside cloud-synced folders.

What is transmitted. The complete manuscript is sent to the configured language model — not selected extracts. Two checks pass the whole body text, and the reference list is sent separately for parsing. Because the tool bounds the reference section at the end of the file, anything placed after the references, including appendices, is transmitted with them. The PDF file itself, the LaTeX source, and the contents of any image are not sent.

Why the whole document. Anonymisation breaches leak through self-citation to precursor work, which appears only in the reference list, and the hidden-text sweep inspects every page’s rendering instructions — render mode, opacity, colour, size — which cannot be sampled.

What the tool cannot see. It does not read images. An identifying detail visible only inside a figure — a screenshot showing a username, an institutional logo, a file path — will not be detected, and figures will need to be checked manually.

Retention by the provider. None, under the zero-data-retention terms described above.

Training: Submission content is not made available for model training. 

What we will report after CHI 2027

Transparency should not end when submissions close.

At the end of this cycle, we intend to report what happened.

Where feasible and compatible with submission confidentiality, this will include:

  • how frequently each tool was used;
  • how often humans agreed and disagreed with its outputs;
  • false-positive and false-negative analyses where ground truth can reasonably be established;
  • differences across methodological and topical categories;
  • cases escalated because humans considered the automated output inappropriate;
  • feedback from ACs, SCs, reviewers, and authors;
  • changes made during the process;
  • and whether each tool will be retained, modified, or discontinued.

The purpose of calling this a pilot should be to learn.

The standard we should hold ourselves to

CHI is a community that studies technology critically.

It would be inconsistent for us to ask authors to examine the limitations, biases, stakeholders, values, and consequences of technological systems while treating the technology used in our own conference infrastructure as beyond scrutiny.

We should hold these tools to a high standard.

That does not mean demanding that they be perfect before they can be useful. We do not demand perfection from individual reviewers, ACs, matching processes, or existing administrative systems either.

It means asking whether each tool improves the overall sociotechnical review process; identifying where it fails; designing human oversight appropriate to the consequences of those failures; looking for disparate effects across our extraordinarily diverse research community; protecting confidential submissions and participant data; and being transparent about what we learn.

Most importantly, it means keeping responsibility where it belongs.

Tools can check. Tools can search. Tools can flag. Tools can suggest. Tools can produce reports.

People make CHI’s scholarly decisions—and people remain accountable for them.