Simplifying administrative language is a key step in improving communication between institutions and citizens. The goal is to make texts clear, understandable and accessible, fostering transparency and inclusion. Because for a public institution, inclusivity is not just compliance with accessibility requirements: it means removing cognitive barriers — and the first cognitive barrier of a public service is almost always its text.
A real-world example of a "hard-to-read" text:
"The recognition and determination of the allowance amount are carried out taking into account the type of family unit, the number of its members, and the total household income. The benefit is granted in decreasing amounts for progressively higher income brackets and ceases upon reaching exclusion thresholds that vary depending on the family type."
The people who need to claim that allowance — often the people who need it most — deserve better.
Before AI: guidelines, metrics, governance
At INPS, simplification was born as a structured process, standing on three legs:
- Plain-language writing guidelines — tone of voice, text structure and formatting, how to avoid bureaucratic language, how to name services;
- An objective measurement framework — the Gulpease readability index, with annual targets;
- Governance — an Experience Authority verifying that the guidelines are applied, with a prudent use of the "silent consent" principle so quality never becomes a bottleneck.
The results of the manual work, before AI even entered the picture:
That move from 57 to 64 looks small, but it crosses the threshold that matters: texts become understandable to people with lower secondary education. In practice, around ten million more people can understand, on their own, what the Institute writes to them.
The limit: manual simplification doesn't scale
Great results — but rewriting content by hand takes a lot of time, and the text production of an institution like INPS is enormous and continuous. There is also a constraint that makes this problem special: simplified texts must retain full legal validity — the semantic content must remain intact. Hence the question that shaped the experiment:
Will AI-simplified texts be able to compete with those produced by human experts?
The experiment: rigor before enthusiasm
We did not just "try ChatGPT on the texts". We built a controlled experiment:
- Multiple large language models compared, to evaluate their effectiveness;
- A dataset of 50 service pages, representative of the different text types, each in three versions: the original bureaucratic one, the version manually simplified by an expert, and the AI-simplified version;
- A rigorous evaluation framework across three dimensions — readability, fluency, content preservation — plus a custom metric, GULBERT, combining the different quality aspects;
- A human in the loop verifying the formal correctness of every generated text;
- And above all: validation with real users.
The analytical results
On readability, AI reaches a score nearly identical to the human experts'. On meaning preservation — measured with BERTScore — the AI version even scores slightly better. But analytical metrics are not enough: there is no guarantee they reflect human judgment.
The acid test: 1,620 users, blind
Innovation is only useful if it brings an actual, perceived benefit. So we ran a large-scale A/B test: 1,620 participants, 5,000 completed evaluations. Each tester read a pair of texts — the original and a rewritten version, alternately produced by the human expert or by the AI — presented in random order and with no indication of authorship. After each reading, a questionnaire on ease of reading, clarity, fluency, confidence to act, tone and style, attention and interest; at the end, a straight preference.
The results:
- Both the expert version and the AI version were clearly and consistently preferred over the original across all 50 services;
- 71.4% of users preferred the AI version over the original;
- The AI version scored above the original on every metric, and consistently equal to or better than the human version — with the largest gains precisely on fluency and on attention and interest.
And the users' own words say more than the numbers:
"In version A my attention dropped by the third line. Long, boring, pedantic. In version B, the structure and the language help you grasp the key points very quickly."
A test participant (blind evaluation)
What we learned
- Generative AI can simplify texts with results equal or superior to human experts, across every indicator we analyzed — provided you keep a human in the loop and strict constraints on content preservation.
- User validation is non-negotiable. Analytical evidence must be confirmed by real people: in our case, 1,600+ users validated not only the AI approach but also the significance of the Gulpease index as a readability measure.
- In the AI hype cycle, it's easy to lose direction. Prioritizing user value is what ensures AI investments lead to real impact. This experiment is an example of virtuous innovation: a new tool applied to a real problem, with measurable value for people.
Next steps
The prototype is becoming an everyday working process, with an interface supporting the human-in-the-loop approach. The next directions: generating texts from scratch starting from primary legal sources (the regulations), rather than from already-written texts; and a source-referencing system linking every part of a text to its corresponding regulation, making legal validation faster and more robust.
Clarity is not an embellishment: for a public service it is the most concrete form of respect — and, as these numbers show, it is now a problem we know how to measure and solve.