czwartek, 8 października 2026

...Whose Bias Counts as Bias? A Word on NewsGuard's Report on Chinese AI Models

The text below is a commentary on NewsGuard’s report on Chinese AI models, which circulated widely in the media in mid-2025. It is not a defense of those models, nor an attempt to downplay the problem of bias in artificial intelligence. It is, rather, a contribution to the debate about how that bias is measured — and about whose criteria we adopt as a benchmark when we declare that a model “fails” 60 percent of the time. The author suggests that figures of this kind say more about the method than about the systems being tested — and that the same caveat is worth applying to Western models. The text does not claim to settle the dispute. It is an invitation to a conversation about how we should measure bias in machines that — like people — have no neutral vantage point.

Reading in Politico, by Anouk Schlung and Larissa Kögl — ‘…A recent NewsGuard investigation found the five leading Chinese‑backed AI models repeatedly failed to correct pro‑China falsehoods.’ ...



I’ll sum it up this way:

NewsGuard has published an audit claiming that five leading Chinese AI models gave inaccurate answers 60 percent of the time when asked about narratives promoted by Beijing. The number is striking. The trouble is that it says more about the method than about the models.

First, “inaccurate” is defined here by editorial judgment. NewsGuard itself decides what counts as a “false narrative” and what does not. That is not a neutral criterion — it is a particular perspective, rooted in a particular political and media tradition. One may agree with it, but one cannot pretend it is a vantage point of objective truth. There is no view from nowhere — not in journalism, and not in evaluating AI systems.

Second, the very design of the study assumes there is one correct answer to contested questions. But in geopolitics, history, or China’s domestic politics, the “correct answer” depends on whom you ask and in what context. A model that echoes Beijing’s official position is not necessarily “misinformed” — it may simply be reproducing the discursive framework it was trained on. Western models do exactly the same thing: they reproduce liberal frameworks and often cannot step outside them on topics where those frameworks are equally one-dimensional.

Third, the report says nothing about how Western models would fare on an equivalent test — one built around topics where the dominant Western narrative is just as resistant to falsification. Had such a test been run, the resulting "fail rate" might well prove to be less a specifically Chinese characteristic than a feature of large language models in general: they reproduce the priorities embedded in their training data and guardrails, not "objective truth."

None of this means Chinese models are impartial. They are not — and no serious person claims otherwise. What it does mean is that the 60 percent figure is not a measure of these models’ “deceptiveness.” It is a measure of the gap between their frameworks and NewsGuard’s. Those are two different things.

The real question is not “how often did the model get it wrong?” but “whose framework does it reproduce, and how consciously can one work with it?” NewsGuard does not ask that question — because asking it would require admitting that its own criterion, too, is somebody’s priority rather than a neutral standard.

Edited with the assistance of DeepSeek AI.

___________________________________________


Tekst poniżej to komentarz do raportu NewsGuard o chińskich modelach AI, który szeroko obiegł media w połowie 2025 roku. Nie jest to obrona żadnego z tych modeli ani próba unmiejszenia problemu stronniczości w systemach sztucznej inteligencji. To raczej głos w sprawie samego sposobu mierzenia tej stronniczości — i pytania, czyje kryteria przyjmujemy za punkt odniesienia, gdy ogłaszamy, że jakiś model „zawodzi” w 60 procentach przypadków. Autor zauważa, że liczby tego typu mówią więcej o metodzie niż o badanych systemach — oraz że to samo zastrzeżenie warto zastosować do modeli zachodnich. Tekst nie rości sobie pretensji do rozstrzygnięcia sporu. Jest zaproszeniem do rozmowy o tym, jak w ogóle mierzyć stronniczość maszyn, które — podobnie jak ludzie — nie mają neutralnego punktu widzenia.

...czytając za "Politico" by Anouk Schlung and Larissa Kögl "...A recent NewsGuard investigation found the five leading Chinese-backed AI models repeatedly failed to correct pro-China falsehoods." skwituję to tak:

NewsGuard opublikował audyt, z którego wynika, że pięć czołowych chińskich modeli AI odpowiadało błędnie w 60% przypadków na pytania o narracje promowane przez Pekin. Liczba robi wrażenie. Problem w tym, że mówi więcej o metodzie badania niż o badanych modelach.

Po pierwsze, „błąd” jest tu zdefiniowany przez redakcyjny osąd. NewsGuard sam ustala, co jest „fałszywą narracją”, a co nie. To nie jest neutralne kryterium — to konkretna perspektywa, zakorzeniona w określonej tradycji politycznej i medialnej. Można się z nią zgadzać, ale nie można udawać, że jest ona punktem odniesienia „obiektywnej prawdy”. Nie istnieje widok znikąd — ani w dziennikarstwie, ani w ocenie systemów AI.

Po drugie, sama konstrukcja badania zakłada, że istnieje jedna właściwa odpowiedź na pytania o sprawy sporne. Tymczasem w geopolityce, historii czy polityce wewnętrznej Chin „właściwa odpowiedź” zależy od tego, kogo pytamy i w jakim kontekście. Model, który powtarza oficjalne stanowisko Pekinu, nie musi być „zdezinformowany” — może odtwarzać ramy dyskursu, w których został wytrenowany. Dokładnie tak samo, jak modele zachodnie odtwarzają ramy dyskursu liberalnego, często nie potrafiąc wyjść poza nie w tematach, w których są one równie jednowymiarowe.

Po trzecie, raport nie mówi nic o tym, jak modele zachodnie wypadłyby w analogicznym teście — skonstruowanym wokół tematów, w których dominująca narracja zachodnia jest równie mało podatna na falsyfikację. Gdyby taki test przeprowadzić, prawdopodobnie okazałoby się, że „fail rate” nie jest cechą specyficznie chińską, tylko właściwością wszystkich dużych modeli językowych: odtwarzają one priorytety swoich danych treningowych i filtrów, a nie „prawdę obiektywną”.

Nie znaczy to, że chińskie modele są bezstronne. Nie są — i nikt rozsądny tego nie twierdzi. Znaczy to natomiast, że liczba 60% nie jest miarą „kłamliwości” tych modeli, tylko miarą rozbieżności między ich ramami a ramami NewsGuard. A to dwie różne rzeczy.

Prawdziwe pytanie nie brzmi: „ile razy model się pomylił?”, tylko: „czyje ramy odtwarza i jak świadomie można z nich korzystać?”. NewsGuard tego pytania nie stawia — bo stawiając je, musiałby przyznać, że jego własne kryterium też jest czyimś priorytetem, a nie neutralnym standardem.