Finding the Right Reviewers: Building a New Paper-Reviewer Matching System for CHI Papers
Authors: Anna Cox, Tony Tang, Thomas Kosch, Erick Oduor, Regan Mandryk
Date: 2026-05-28
A central goal of the CHI papers review process is to find the right reviewers for each paper. By “right,” we mean reviewers who meet the minimum qualifications for reviewers and who have the appropriate expertise to review the paper – both topically and methodologically. A good reviewer match is not just someone who knows the broad area of a paper. It is someone who can understand what the paper is trying to do, evaluate it in relation to the right scholarly traditions, and provide feedback that is fair, constructive, and useful.
As CHI continues to grow in size and breadth, this matching problem becomes more important. CHI now brings together many HCI communities, methods, topics, contribution types, and epistemological traditions. This diversity is one of CHI’s great strengths. It is also a challenge for peer review: a paper on a health intervention, a critical design artefact, an autoethnography, a computational model, a fabrication technique, and a theory paper may all be “CHI papers,” but they require different kinds of reviewing expertise.
In the rest of this blogpost we answer:
- How have we been matching reviewers to papers up til now?
- What are the problems with that process?
- How will new descriptors help?
- How will the new descriptors be used?
- How did the proposed descriptor structure emerge? (only for the very curious!)
- What principles have guided the work?
- What will authors and reviewers see?
- Who has helped to develop this?
- What do we need from the community? (PLEASE read this!)
- What happens next?
1. How have we been matching reviewers to papers up til now?
Authors submitted papers to subcommittees. Associate Chairs (ACs) reviewed titles and abstracts and bid on papers they would like to handle. Once the papers had been allocated to ACs, ACs were required to find reviewers. Whilst in many cases reviewers used their own networks to find good reviewers, reviewer assignment in PCS were also supported by a recommendation tool that produced match scores between submitted papers and potential reviewers. These scores drew on two main signals. First, PCS compared the keywords selected for the paper with the expertise-weighted keywords in each reviewer’s profile: in simplified terms, the score reflected the overlap between the paper’s keywords and the reviewer’s stated areas of expertise, normalized against an “ideal” reviewer who is an expert in all of the paper’s keywords. Second, PCS can use sample papers uploaded by reviewers, comparing those documents with the submitted paper using latent semantic analysis, following the approach described by Charlin and Zemel in their 2013 paper “The Toronto Paper Matching System: An automated paper-reviewer assignment system“. In practice, PCS extracts word counts from the documents and used them to estimate topical similarity. This means that the quality of reviewer recommendations depended heavily on the quality, completeness, and specificity of the keywords and expertise information available in the system.
2. What are the problems with that process?
The current keyword system was not designed to carry so much responsibility. In practice, reviewer profiles are often incomplete or out of date: many people set up their PCS profiles years ago and have little reason to revisit them unless explicitly prompted. The system also relies heavily on self-declared expertise, which can be noisy. Some reviewers may select very broad areas, or indicate expertise across too many topics, making it harder to distinguish deep expertise from general familiarity. Over time, the keyword list itself has also tended to grow by accumulation rather than revision: new terms have been added, but older, overlapping, vague, or low-signal terms are rarely removed or restructured. As a result, many keywords do not meaningfully differentiate submissions or reviewers. Broad labels such as “interaction design,” “user experience,” or “empirical study” may describe large parts of CHI, but they do little to identify who is best placed to evaluate a specific paper. Fundamentally, this means the existing keyword system is a weak starting point for matching: it asks too much of terms that are often stale, overly broad, inconsistently used, or insufficiently tied to the actual expertise needed for fair review.
Ultimately, the descriptors approach is trying to ensure better routing for submissions to the reviewer teams best equipped to review the work.
- “They didn’t get it.” – As authors, it is frustrating when our work is reviewed by someone with a different epistemological framing than the intended audience. The descriptor approach helps to identify reviewers that understand a submission from the right epistemological framing.
- Feeling stretched as reviewers. As reviewers, it is easier to review works that are (currently) interesting to us where we are familiar with the current literature. When we are asked to review work based on older expertise, or that is just outside is just outside this is harder to review.
- Finding reviewers. Our matching system is currently easily gamed/made inefficient by profiles that indicate “expert” expertise on every keyword. Resetting this and limiting what constitutes “expertise” will help for better routing.
- Many current keywords are no longer useful for routing. “Qualitative Methods” (23.6%), “Quantitative methods” (17.0%), and “User Experience Design” (11.6%) were some of the most popular tags at CHI 2026. While these are helpful for authors explaining what their work is, they are not quite as useful to identify reviewers. On the other hand, a term such as “Sustainability” was tagged by 2% of submissions. Such a keyword is more targeted, and if a reviewer were to signal that as an area of competency, allows us to route a submission more effectively.
3. How will new descriptors help?
The descriptors we use to describe our expertise may sound like a small piece of conference infrastructure. They are not. They are the key to matching reviewers to papers. These descriptors are one of the ways that authors tell us what their work is about. They are also one of the ways that reviewers could tell us what they are able to review. If the system is too vague, too broad, too inconsistent, or too tied to historical categories that no longer route papers well, then reviewer matching suffers.
Good descriptor words help us answer questions such as:
- What is the paper about?
- What methods, approaches, or epistemological traditions does it use?
- Who is the work with, for, or about?
- What kind of contribution is the paper making?
- Which reviewers are likely to understand both the topic and the way the paper argues for its contribution?
This matters because we need CHI to move toward a more scalable review model in which assignments balance expertise, conflicts of interest, and workload distribution. In the proposed sustainable and scalable review model, reviewer assignment depends on expertise profiles and matching scores, with ACs using those profiles to identify appropriate reviewers from the volunteer pool (i.e., those who have submitted to CHI that year) and externally.
Historically, we have relied on the keywords that authors use to describe their paper to potential readers (and future systematic review searches). These keywords tend to describe the content of a paper, and might not be best suited for finding the right expertise to assess a submitted manuscript. We need up-to-date competency or expertise descriptors—hereafter called descriptors—that characterise a person’s reviewing expertise and what expertise is needed to evaluate a submitted paper.
4. How will the new descriptors be used?
The descriptor system is intended to support reviewer matching. It is not intended to rank papers, judge quality, define what “counts” as CHI, or replace human judgment.
In the full review process, good matching means helping ACs identify reviewer teams that collectively cover the relevant topic, method, and contribution expertise. This aligns with the broader restructuring goal of supporting fair, sustainable, and scalable peer review at CHI. Our intention is that descriptors will be used at multiple stages:
- For routing papers to the right reviewers
- Helping authors to indicate the reviewing competencies for their submission
- Helping ACs during bidding (surfacing submissions that are more closely aligned with their stated expertise)
- Helping ACs to identify external reviewers with the appropriate competencies
- Helping SCs with assignments of papers to ACs with the appropriate competencies
- For reviewer/committee recruitment and forecasting
- Helping ACs to understand whether the competencies of the recruited external reviewers team covers the competencies requested by authors
- Helping the Technical Program Chairs (TPCs) to understand whether the competencies of the Paper Chairs as a whole covers the needed competencies of the submissions (and, trend-wise, where needs will be in the near future)
Procedurally: For Authors
At submission time, authors will indicate descriptors for their submission. They should think of these descriptors as being primarily answering the question: “What competencies within the CHI community are necessary to review this submission properly?”
Procedurally: For Reviewers and Committee Members
Reviewers and Committee Members provide their expertise based on these descriptors.Our intention is to use these expertise as markers that indicate, “Have recent publishing experience, and am interested in reviewing.” While many members of our community have broad expertise, we are asking each person to signal a smaller, focused subset of this expertise. This allows for more effective routing.
In phases of the review process where algorithmic matching may be used, such as proposed rapid triage or assisted desk rejection (ADR) processes or presenting a subset of papers to ACs for bidding, the quality of matching becomes even more important. The ADR proposal explicitly notes that CHI’s heterogeneity creates risks when reviewers are poorly matched to methods or domains outside their expertise, and that good matching and multiple ratings are needed to reduce these risks.
The descriptor system will therefore be used as one signal among others, alongside paper text, conflicts of interest, reviewer workload, reviewer qualifications, and human oversight. It is a tool to support better decisions, not to automate judgment.
5. How did the proposed descriptor structure emerge?
**Only read if you are curious in the process of how these descriptors were developed.**
The Paper Chairs for CHI2027 have conducted a number of analyses of the papers that were submitted to CHI2026. We investigated a number of different methods for extracting descriptors that included:
- The keywords selected by the authors in PCS for their paper
- The author keywords provided in the paper
- The titles and abstracts of papers
- The full paper text
- The subcommittee that the authors chose to submit to.
The proposed structure emerged through a combination of computational assistance and manual curation.
We began with author keywords from recent CHI submissions: 5,379 unique author keywords across 6,730 submissions. These were normalized and clustered into approximately 4,632 canonical leaf terms.
From there, we created a seed taxonomy: a structured file containing proposed sub-pile names and seed phrases. We then used embedding similarity to assign each canonical author keyword to the closest seed category. This produced a first-cut board in which every leaf term had a home.
That first version was not treated as final. It was a working object for curation. We used a repeated loop:
embed → assign → read → restructure → re-embed
In each round, we inspected the board, looked for categories that were too broad, too narrow, redundant, ambiguous, or misplaced, edited the seed taxonomy, and reclassified the remaining terms. Accepted leaves stayed where they were; uncertain or poorly matched terms moved as the structure improved.
During this process, we realized that another important signal was already present in the historical submission data: the subcommittee to which authors had submitted their work. We therefore incorporated subcommittee information into the analysis as an additional signal. This helped us understand how author keywords related to CHI’s existing review structure, while also making visible places where historical subcommittee labels were not sufficiently precise for future matching.
Once we had a structure that seemed workable, we built two prototype tools:
- An author-facing tool that shows what descriptors an author might select when submitting a paper.
- A reviewer-facing tool that lets reviewers indicate their areas of reviewing expertise.
We have tested these tools ourselves to see whether they make sense for papers and expertise we know well. Now we need the CHI community’s help to test that they make sense for everyone.
6. What principles have guided the work?
Several principles have shaped the curation of the new descriptor system.
First, we are treating the descriptor list as review infrastructure. It is not simply a list of topics. It is part of how papers and reviewers will be routed to one another.
Second, we want to honour existing CHI knowledge. The current and historical subcommittees are not arbitrary; they reflect years of community practice and encode a great deal of knowledge about how CHI has understood its areas. In that sense, subcommittees function as “big keywords.” At the same time, prior analysis of CHI’s program committee history shows that subcommittee-based structures can also create unevenness and silos, with different subcommittees developing different reviewing cultures and expectations over time. Our aim is to preserve useful domain knowledge without reproducing unnecessary boundaries.
Third, we are treating author keywords as a form of ground truth. Authors have already told CHI what their work is about through the keywords they chose at submission. Rather than invent a taxonomy from scratch, we began from those author-provided terms.
Fourth, we are trying to make the facets as orthogonal as possible. Domain, method, users, and contribution should not collapse into one another. “Health” is not a method. “Interview study” is not a domain. “Older adults” is not a contribution type. Keeping these signals separate should help us match papers to reviewers more accurately.
Fifth, we are treating the sub-pile name as the unit of meaning. In practice, this means that the named category should be meaningful enough for authors and reviewers to recognize, while the author keywords beneath it act as more specific leaves.
Sixth, we are being careful with legacy terms. Some familiar terms, such as “interaction design” or “UX design,” may be meaningful in some contexts, but can also be too broad to route papers effectively. A descriptor that almost anyone can select may not help us find the right reviewer.
Finally, the Method / Approach facet is intended to capture more than technique. It should also capture stance: for example, whether a paper is doing controlled experimentation, interpretivist qualitative work, research through design, critical scholarship, computational modelling, system-building, participatory research, or another approach. This matters because two papers in the same domain may need very different reviewers if they make knowledge in different ways.
7. What will authors and reviewers see?
Our current design asks authors and reviewers to describe papers and expertise using a small number of facets:
Domain — what the work is about.
For example: Health & Wellbeing, Accessibility, Learning, Sustainability, Games and Play.
Method / Approach — how the work is done, including epistemological stance.
For example: qualitative research, research through design, critical approaches, computational modelling, systems building, quantitative studies, theory development.
Users / Participants / Communities — who the work is with, for, or about, where applicable.
For example: children, older adults, healthcare workers, disabled people, creators, workers, families, specific communities, or no specific user group.
Primary Contribution — what kind of contribution the paper makes.
For example: empirical finding, system, theory, dataset, method, design artefact, infrastructure, critique.
This structure is intended to help separate signals that are often mixed together. A paper can be in the domain of health, use a qualitative or computational approach, involve clinicians or patients, and contribute a system, empirical insight, or theory. These are different kinds of information, and we want the descriptor system to preserve those distinctions.
8. Who has helped to develop this?
The Paper Chairs have used data from the past few years of submissions. We have discussed this with the TPCs and GCs for CHI2027. The SCs from CHI2026 and conference leaders from other SIGCHI conferences have been invited to review this and have helped us to find additional descriptors and to refine the piles.
9. What do we need from the community?
We are now asking members of the CHI community (you!) to help us test the proposed descriptor structure and the prototype tools.
In particular, we want to know:
- Are the domains complete?
- Are any domains missing, duplicated, or named in a way that excludes part of the community?
- Do the Method / Approach categories capture the range of ways CHI papers make knowledge?
- Are contribution types clear and useful?
- Are user, participant, and community categories appropriately scoped?
- Are there terms that feel too broad to be useful for reviewer matching?
- Are there terms that may reinforce outdated assumptions about areas, methods, communities, or kinds of contribution?
- When you describe your own work, do the available choices let you do so accurately?
- When you describe your reviewing expertise, do the choices help you say both what you can review and what you should not be asked to review?
We are especially interested in feedback from communities whose work may not have been well represented in past taxonomies, whose methods are often misunderstood, or whose contributions cut across established CHI boundaries. A keyword system can only support fair matching if it reflects the breadth of the community it is meant to serve.
Note that the descriptor tools we are sharing here are just for testing out the piles. At this stage, we don’t expect that these tools will be part of the submission system, so please don’t give feedback on the UX! 🙂
10. What happens next?
We will use community feedback to revise the domains, methods, user categories, contribution types, and underlying keyword mappings. We expect this to remain iterative. CHI changes, HCI changes, and the vocabulary of our community changes with it.
Our goal is not to create a perfect taxonomy. No fixed list can fully capture the complexity of CHI. Our goal is to create a better shared infrastructure for matching papers to reviewers: one that is transparent, revisable, and grounded in the language authors and reviewers actually use.
CHI’s values call on us to be transparent, inclusive, accessible, and bold in adapting to the changing needs of the community. Building this descriptor system in conversation with the community is one way to put those values into practice.
We invite you to try the tools, describe your own work and expertise, and tell us where the structure works, where it fails, and what it is missing. Your feedback will help us build a review process that better supports the breadth, quality, and diversity of CHI research.
We will of course be testing the matching algorithm in multiple ways to ensure that it works effectively.
Please try out the tools:
- Go to https://hcitang.github.io/apps/chi2026-keyword-feedback/ and try to communicate the expertise needed to review a couple of your papers and explore describing your own expertise as a reviewer.
- Go to https://forms.gle/K3khuGurmWwKgLJn6 to send us feedback.