Home / Research / Research Brief: AI summaries of staff feedback drop concerns raised once

Research Brief: AI summaries of staff feedback drop concerns raised once

An audit of 2,586 staff comments finds LLM summaries keep criticism but drop concerns raised once, 86% of the time. What to ask of your feedback tools.

Research · Luminary Research Brief · 2 October 2026 · 4 min read

If you use an AI tool to summarise staff comments for managers, a new preprint audit suggests it may be quietly discarding the rarest concerns. Tamme, Hantel and Khosrawi-Rad found that a large language model (LLM) summary pipeline kept criticism but dropped a concern voiced once 86% of the time. The filter worked on how often something was said, not on whether it was positive or negative. For an operator, that means the summary may repeat what everyone already knows and lose the one warning that matters.

What the researchers did

The authors worked with 2,586 free-text responses from a global professional service company. The responses were in English and German and came from a weekly team feedback exercise with two prompts: one asking what went well (appreciation) and one asking what to improve.

They then looked at 45 summaries written for leaders by an LLM pipeline. For the themes in the raw comments, they checked whether each theme made it into the summary. They call the resulting measure a Voice Retention / Representation Ratio. In plain terms, it asks what share of a given group of voices survives summarisation. They split the results by how many people raised a theme, how long the comment was, and which language it was written in. They also tested whether a prompt aimed at specific themes could improve retention.

I only had the paper’s abstract, introduction and conclusion to work from, not the methods or results sections. The detail below is limited accordingly.

What they found

Employees withheld praise far more than criticism. Respondents answered the improvement prompt 99.8% of the time but the appreciation prompt only 80.9% of the time. The authors put withholding praise at roughly 82 times more common than withholding criticism. Much workplace research worries about silence on problems, but here the gap fell on the positive side. The input reaching the summariser was already tilted towards change.

The summary filtered by popularity. Criticism survived almost everywhere. What did not survive was infrequent voice: a concern raised once was dropped 86% of the time. The abstract reports theme retention of 0.14 against 0.74. The 0.14 matches the 86% drop for once-voiced concerns. The abstract does not spell out the comparison group for 0.74, so I would not read more into it than that.

Short and German-only comments were lost along the same axis. The authors describe the German result as directional, meaning suggestive rather than firmly established.

Sentiment made no independent difference. Once the authors controlled for how often a theme appeared, whether it was positive or negative had no separate effect. They conclude the harm is driven by prevalence. This matters because an audit that only checks whether a summary is too positive or too negative would miss it.

A targeted prompt helped only in part. It recovered the themes it was told to look for. The authors say it left the frequency and length bias in place, so prompt-level fixes are partial.

What this means for your business

The paper does not test remedies beyond the targeted prompt, so the actions below are my reading of the findings, not tested advice.

  • Treat the summary as a list of common themes. It will tell you what is said most often. It is a poor guide to what is said once, and the paper suggests it will usually be silent on that.
  • Keep the raw comments within reach. If managers only ever see the summary, nobody can notice what was cut. Have someone read a sample of the original comments each cycle, with particular attention to short and one-off ones.
  • Name the themes you cannot afford to miss. Safety, harassment, fraud and client risk are examples. A targeted prompt recovered named themes in the study. It will not catch a problem you did not think to name.
  • Ask vendors or your own team for retention figures. The authors propose a disaggregated voice-retention card, which reports retention by frequency, length and language. If a tool cannot show how often a once-mentioned concern survives, you do not know its blind spot.
  • Check multilingual teams separately. If staff write in more than one language, test each one. The German finding is only directional, but it is cheap to check.
  • Do not rely on a sentiment check alone. A summary can look balanced in tone and still drop the minority view.

This matters most in smaller firms. With a few dozen comments, a single concern is often a large share of the signal, and a quick read of the raw text costs little.

Limits worth knowing

  • One company, one sector. The data come from a single global professional service firm, in English and German. Retail, manufacturing or a ten-person agency may behave differently.
  • A small number of summaries. The retention results rest on 45 leader-summaries, and the German effect is flagged by the authors themselves as directional.
  • One kind of pipeline. The results describe the pipeline the authors tested. Another model, prompt or summary length could retain more or less. The paper’s own mitigation test suggests the pattern is not easily prompted away, but that is one test.
  • Preprint. The paper has not been peer reviewed.
  • Partial reading. I did not see the full methods and results, so I cannot judge how themes were coded or how the models were configured.

On balance, the central finding deserves moderate weight. It is a field audit on real data rather than a lab exercise, and its mechanism (summaries favour what is repeated) is plausible. But it is one firm and one pipeline. Use it as a reason to test your own tools, not as proof of how every summariser behaves.

Source: Thilo Tamme, Anton Hantel, Bijan Khosrawi-Rad, “Whose Voice Survives the Summary? A Voice-Retention Audit of LLM Employee Listening”, arXiv. This paper is a preprint and has not yet been peer reviewed.

How Luminary Solutions approaches this

At Luminary Solutions, we turn findings like this into working systems for small and mid-sized businesses: the processes, automations and checks that put the research to use. If this brief raised a question about your own operation, let’s talk.

Explore how we work →

LM
Luminary Media Editorial
The Luminary Research Brief translates one new academic paper each week into practical insight for founders and operators.

Stay ahead with Luminary Media

Weekly insights on AI automation, marketing systems and digital strategy, delivered to your inbox.

Subscribe now →

You Might Also Like