writing · llm research
frictions/attriti
Recurring ways language models fail, named and countered, from inside an extended writing and research project run as a standing testbed.I modi ricorrenti in cui i modelli linguistici cedono, nominati e contrastati, dall'interno di un esteso progetto di scrittura e ricerca condotto come banco di prova continuativo.
A model rarely fails loudly. It agrees, it returns something plausible, and the work drifts — three sessions later you are re-arguing a term that was settled weeks ago. Benchmarks cannot see this. On month three of a project, it is most of what you see.
It is the reason these phenomena are visible at all: most of them need weeks of accumulated context before they surface, which is exactly the condition an afternoon of testing cannot reproduce. It is also the limit of the evidence. This is one long case observed systematically, not a statistic, and nothing here is presented as a measurement.
Most of this work was done with Claude, and Claude is named wherever the episode is Claude's. Where the literature documents the same behaviour across vendors — positional degradation over long inputs, sycophancy under preference training, the limits of unaided self-correction — it is described as a family trait rather than one company's fault.
This inventory was produced by applying the protocols it describes. The observations, the analysis, the protocols and the final revision are mine; the drafting was done in dialogue with the systems under observation, under those same protocols — blind fields, forced refutations, pre-declared failure criteria.
articlessix
italianoenglish
01
italiano
I limiti della mia chat sono i limiti del mio mondo
Che cosa sopravvive alla fine di una sessione, che cosa no, e il documento di trasmigrazione che porta un progetto attraverso lo stacco.
english
The Limits of My Chat Are the Limits of My World
What survives the end of a session, what does not, and the transmigration document that carries a project across the gap.
02
italiano
Il modello vuole darsi ragione
Chiedi un test e ottieni un test che conferma l'ipotesi. Il campo cieco: come si progetta una verifica contro il modello anziché con lui.
english
The Model Grades Its Own Homework
Ask for a test and you get a test that confirms the hypothesis. The blind field: how to design verification against the model rather than with it.
03
italiano
Ipse dixit: dalla sicurezza al dogma
Il guasto comincia quando sparisce il margine di incertezza. Quattro segnali precoci, e perché interrompere la catena funziona meglio che correggerla.
english
Ipse Dixit: From Confidence to Dogma
The failure starts when the margin of uncertainty disappears. Four early signals, and why breaking the chain beats correcting it.
04
italiano
«Riprova» è il prompt peggiore
Gli errori di un modello non incontrano mai un rimborso in token — ma i token non sono mai stati la parte cara. Prompt diagnostici per il secondo fraintendimento.
english
"Try Again" Is the Worst Prompt
There are no token refunds for a model's mistakes — but tokens were never the expensive part. Diagnostic prompts for the second misunderstanding.
06
italiano
Complessità ingenua: il salto logico invisibile
La risposta arriva col ragionamento già speso. Perché è costoso da rettificare e pericoloso da accettare.
english
Naïve Complexity: The Invisible Logical Leap
The answer arrives with its reasoning already spent. Why that is expensive to correct and dangerous to trust.
field notesone
italianoenglish
05
italiano
Il collasso verso il tipico
Chiedi cinque alternative e ottieni cinque versioni della stessa cosa. Tre condizioni su quattro modelli, un repertorio condiviso, e una cosa che non si muove: il gesto.
nota operativa
english
The Collapse Toward the Typical
Ask for five alternatives and get five versions of one thing. Three conditions across four models, a shared repertoire, and one thing that does not move: the gesture.
field note
case reportsone
italianoenglish
08
italiano
Nove contromisure, testate ripetutamente per due settimane
Nove interventi effettuati nell’arco di due settimane. Una linea separa gli interventi riusciti dagli insuccessi.
caso studio
english
Nine Countermeasures, tested repeatedly for two weeks
Nine interventions applied across a two-week research phase, each with its outcome. One line separates what held from what did not.
case report
Every article closes on a question its protocol does not answer.
This page opens on the one they share: how much of what a model gets wrong is available to be noticed?
diego ballarin — editorial writer and researcher
emaildiegoballarin@icloud.com · phone number+39 3479280321 · Instagram@diegoballarin · LinkedIn/in/diego-ballarin
emaildiegoballarin@icloud.com · phone number+39 3479280321 · Instagram@diegoballarin · LinkedIn/in/diego-ballarin
© Diego Ballarin · CC BY-NC 4.0
github.com/diegoballarin-reason/frictions