sobota, 15 sierpnia 2026

Two Hundred Tokens, Three Little Icons, and One Human Being

A weekend column on how the European Union decided to teach us how to recognise artificial intelligence



Sometimes one gets the impression that the world is complicated.

Fortunately, we have the European Union.

The Union has a remarkable ability to make things complicated in a manner that is orderly, documented, and accompanied by an appropriate icon.

Recently, I took an interest in the European rules concerning transparency of AI-generated content. At first, the matter seemed quite straightforward.

If AI generates something, it should be labelled.

If a human writes something, it need not be.

Simple?

Not quite.

It turns out that there are AI-generated texts, AI-modified texts, texts involving AI, texts reviewed by humans, artistic texts, satirical texts, fictional texts, and texts concerning matters of public interest.

And all of them require rather different treatment.

Fortunately, there are icons.

Three of them.

The first indicates content generated entirely by AI.

The second indicates content created by a human and partially modified by AI.

The third is the basic one, indicating AI involvement in general.

This is reassuring.

The modern citizen will therefore be able to visit a website, look at a piece of content, and immediately know what he or she is looking at.

Provided, of course, that the citizen has first learned to recognise the three official European icons.

But matters become even more interesting.

The technical part of the arrangement provides for watermarking — a digital watermark. And here we encounter something that particularly caught my attention.

For free-form text exceeding 200 tokens, watermarking remains required, while its reliability for shorter texts may be lower.

So we already have our first threshold.

Two hundred tokens.

I have no idea why 200.

Perhaps 199 tokens are not quite enough for a human being to realise that a machine has deceived him.

Perhaps, at 201, the matter becomes sufficiently serious to require a watermark.

I shall not ask.

I should not wish to provoke the creation of another working group.

For what, after all, is a token?

Is one token one word?

Not necessarily.

Is punctuation a token?

It depends.

Is “AI” one token?

Possibly.

Or perhaps two.

And so we have arrived at the frontier of a great European question:

Is a text containing 199 tokens a short text, while one containing 201 tokens is a long one?

And if so, what precisely is a text containing 200 tokens?

That question is already serious enough to justify the appointment of a European Expert for Borderline Tokens.

Then, naturally, a committee.

Then a subcommittee.

Then an explanatory document explaining that 200 tokens means 200 tokens, except where, in a particular context, it means something else.

And having settled that, we may proceed calmly to the next problem.

What, for instance, are we to do with art?

Here, I must admit, the regulator has shown a commendable degree of common sense.

If a deepfake forms part of an artistic, satirical, fictional or similar work, the disclosure need not necessarily be placed directly in the work if doing so would spoil its presentation. It may instead be placed beside it, in accompanying material, or in introductory credits.

Thus, if a painter creates a picture with the assistance of AI, there is no need, apparently, to paint AI across the middle of the canvas.

One may put it beside the picture.

Perfectly sensible.

I do, however, have certain reservations concerning satire.

If a satirist produces a satire with the assistance of AI and then attaches an official European icon to it, the reader may have some difficulty deciding which part is the satire and which part is the European system of labelling.

Fortunately, when it comes to text concerning matters of public interest, there is another solution.

A human being.

If the text has undergone human or editorial review, and a natural or legal person assumes editorial responsibility for publication, the disclosure obligation may be waived entirely.

And here I stopped for a moment.

Because suddenly, after all the technology — watermarks, metadata, icons, detection systems, interoperability, tokens and the rest — the solution to the problem may simply be:

A human being has read it and takes responsibility.

How terribly old-fashioned.

A human reads.

A human judges.

A human decides.

A human publishes.

A human answers for it.

One almost forgets that this was once how responsibility for the written word worked.

Naturally, the European Union does not intend to leave this human being entirely unattended.

If he is not a professional media provider, he may need a formal policy setting out, among other things, who bears editorial responsibility, what measures and resources are used for pre-publication review and — as the HLC analysis explains — appropriate contact details must be publicly available.

So the human being remains.

He has merely been properly documented.

Fortunately, there are also alternative ways of complying with the requirements.

One may do things differently.

One merely has to explain why.

And the competent market-surveillance authority may then assess the alternative arrangement on a case-by-case basis.

This, I think, is a particularly beautiful expression of regulated freedom.

You may choose.

You may devise your own solution.

You may depart from the standard solution.

Provided that, later, the appropriate authority agrees that you were right.

And so we arrive at the end of our brief excursion into the European world of artificial intelligence.

We have a watermark.

We have metadata.

We have three icons.

We have fully AI-generated content.

We have partially AI-modified content.

We have artistic work.

We have satire.

We have public-interest text.

We have 200 tokens.

We have human review.

We have editorial responsibility.

We have procedures.

We have documentation.

We have alternative measures.

We have a supervisory authority capable of deciding whether the alternative measure was acceptable.

There is only one thing missing.

The text.

But I expect we shall manage that somehow.

Especially since this particular text has been read by a human being.

And that human being — somewhat to his own surprise — still believes that he should decide what he publishes.

For that reason, I shall refrain from displaying any of the three official European icons here.

Not because I have anything against artificial intelligence.

Quite the contrary.

I simply would not wish anyone to confuse the information that AI was used with the rather more important information concerning who is responsible for what you have just read.

These, as it turns out, are two entirely different things.

And perhaps that is why we needed 200 tokens, three icons, a watermark, metadata, procedures, committees of experts and an entire European code of practice to arrive at a conclusion which is considerably older than artificial intelligence:

In the end, a human being should be responsible for the word.

__________________________________________


Dwieście tokenów, trzy ikonki i jeden człowiek

Weekendowy felieton o tym, jak Unia Europejska postanowiła nauczyć nas rozpoznawać sztuczną inteligencję

Czasami człowiekowi wydaje się, że świat jest skomplikowany.

Na szczęście istnieje Unia Europejska.

Unia potrafi bowiem skomplikować go w sposób uporządkowany, opisany i opatrzony odpowiednią ikoną.

Ostatnio zainteresowałem się unijnymi zasadami przejrzystości treści generowanych przez sztuczną inteligencję. I muszę przyznać, że początkowo sprawa wydawała mi się prosta.

Jeżeli AI coś wygenerowała, trzeba to oznaczyć.

Jeżeli człowiek coś napisał, nie trzeba.

Proste?

Otóż nie.

Okazuje się bowiem, że mamy tekst wygenerowany przez AI, tekst zmodyfikowany przez AI, tekst z udziałem AI, tekst poddany kontroli człowieka, tekst artystyczny, satyryczny, fikcyjny oraz tekst dotyczący spraw interesu publicznego.

A wszystko to wymaga odpowiedniego potraktowania.

Na szczęście przygotowano odpowiednie ikonki.

Są trzy.

Pierwsza oznacza treść całkowicie wygenerowaną przez AI.

Druga — treść stworzoną przez człowieka i częściowo zmodyfikowaną przez AI.

Trzecia jest podstawowa i informuje ogólnie o udziale AI.

To bardzo dobrze.

Człowiek współczesny będzie więc mógł kiedyś otworzyć stronę internetową, spojrzeć na materiał i od razu wiedzieć, z czym ma do czynienia.

Oczywiście pod warunkiem, że nauczy się rozpoznawać trzy unijne ikonki.

Ale jest jeszcze ciekawiej.

Techniczna część regulacji przewiduje bowiem watermarking, czyli znak wodny. I tutaj pojawia się rzecz, która szczególnie mnie zainteresowała.

Dla tekstu swobodnego przekraczającego 200 tokenów watermarking pozostaje wymagany, przy czym autorzy opracowania uczciwie przyznają, że przy krótszych tekstach jego niezawodność może być mniejsza. (www.hlc.com)

A więc mamy już pierwszy próg.

200 tokenów.

Nie wiem, dlaczego akurat 200.

Być może 199 tokenów jest jeszcze zbyt mało, żeby człowiek zdążył się zorientować, że został oszukany przez maszynę.

Być może przy 201 tokenach sytuacja staje się już na tyle poważna, że potrzebny jest znak wodny.

Nie będę jednak pytał.

Nie chcę prowokować powstania kolejnej grupy roboczej.

Bo przecież czym jest token?

Czy jeden token jest słowem?

Nie zawsze.

Czy znak interpunkcyjny jest tokenem?

Zależy.

Czy „AI” jest jednym tokenem?

Być może.

A może dwoma?

I oto właśnie znaleźliśmy się na granicy wielkiego europejskiego problemu.

Czy tekst liczący 199 tokenów jest tekstem krótkim, a tekst liczący 201 tokenów tekstem długim?

A jeżeli tak, to co z tekstem liczącym dokładnie 200?

To pytanie jest już na tyle poważne, że należałoby chyba powołać europejskiego eksperta ds. tokenów granicznych.

A potem komisję.

A potem podkomisję.

A potem dokument wyjaśniający, że 200 tokenów oznacza 200 tokenów, chyba że w określonym kontekście oznacza inaczej.

I wtedy moglibyśmy spokojnie przejść do kolejnego problemu.

Co zrobić z dziełem artystycznym?

Tutaj regulator wykazał się zresztą zdrowym rozsądkiem.

Jeżeli deepfake jest częścią dzieła artystycznego, satyrycznego, fikcyjnego lub podobnego, oznaczenie nie musi być umieszczone bezpośrednio w dziele, jeżeli psułoby jego odbiór. Można je umieścić obok, w materiałach towarzyszących albo w napisach początkowych. (www.hlc.com)

Czyli jeśli malarz stworzy obraz przy pomocy AI, nie trzeba koniecznie namalować na środku płótna:

AI

Można napisać to obok.

To rzeczywiście rozsądne.

Mam jednak pewne obawy dotyczące satyry.

Jeżeli satyryk stworzy satyrę przy pomocy AI, a następnie oznaczy ją oficjalną unijną ikonką, odbiorca może mieć problem z ustaleniem, co właściwie jest częścią satyry, a co już jest elementem unijnego systemu oznakowania.

Na szczęście w przypadku tekstu dotyczącego spraw interesu publicznego pojawia się jeszcze jedno rozwiązanie.

Człowiek.

Jeżeli tekst przeszedł kontrolę człowieka lub kontrolę redakcyjną, a konkretna osoba fizyczna albo prawna ponosi odpowiedzialność redakcyjną za publikację, obowiązek ujawnienia wykorzystania AI może zostać całkowicie uchylony. (www.hlc.com)

I tutaj, przyznam, zatrzymałem się na dłużej.

Bo nagle okazuje się, że po całej tej armii technologii — watermarkach, metadanych, ikonach, detekcji, interoperacyjności i tokenach — rozwiązaniem problemu może być zwyczajnie:

człowiek przeczytał i odpowiada.

Jakież to staroświeckie.

Człowiek czyta.

Człowiek ocenia.

Człowiek decyduje.

Człowiek publikuje.

Człowiek odpowiada.

Niemal zapomnieliśmy, że kiedyś tak właśnie działała odpowiedzialność za słowo.

Oczywiście Unia nie pozostawia tego człowieka całkowicie bez opieki.

Jeżeli nie jest on profesjonalnym nadawcą medialnym, powinien mieć formalną politykę określającą między innymi, kto odpowiada redakcyjnie, jakie środki i zasoby zapewniają kontrolę przed publikacją oraz — jak czytamy w analizie HLC — odpowiednie dane kontaktowe mają być publicznie dostępne. (www.hlc.com)

A więc człowiek pozostał.

Tylko że został odpowiednio opisany.

Na szczęście istnieją jeszcze alternatywne sposoby spełnienia wymogów.

Tyle że trzeba je uzasadnić.

A właściwy organ nadzoru może je następnie ocenić indywidualnie.

I tutaj dochodzimy do rozwiązania, które wydaje mi się szczególnie piękne.

Można zrobić inaczej.

Ale trzeba udowodnić, że zrobiono inaczej właściwie.

To jest prawdziwa wolność regulowana.

Możesz wybrać.

Możesz zastosować własne rozwiązanie.

Możesz odstąpić od rozwiązania standardowego.

Pod warunkiem że później odpowiedni organ uzna, że miałeś rację.

I tak oto dochodzimy do końca naszej krótkiej wycieczki po europejskim świecie sztucznej inteligencji.

Mamy watermark.

Mamy metadane.

Mamy trzy ikonki.

Mamy tekst całkowicie wygenerowany.

Mamy tekst częściowo zmodyfikowany.

Mamy tekst artystyczny.

Mamy satyrę.

Mamy tekst dotyczący interesu publicznego.

Mamy 200 tokenów.

Mamy człowieka dokonującego kontroli.

Mamy osobę odpowiedzialną za publikację.

Mamy procedurę.

Mamy dokumentację.

Mamy możliwość zastosowania rozwiązania alternatywnego.

Mamy organ nadzoru, który może to rozwiązanie ocenić.

Brakuje nam już tylko jednego.

Tekstu.

Ale z tym chyba sobie poradzimy.

Zwłaszcza że niniejszy tekst został przeczytany przez człowieka.

I człowiek ten — ku własnemu zdziwieniu — nadal uważa, że to on powinien decydować, co publikuje.

Dlatego na wszelki wypadek nie zamieszczę tutaj żadnej z trzech unijnych ikonek.

Nie dlatego, że jestem przeciwko sztucznej inteligencji.

Wręcz przeciwnie.

Po prostu nie chciałbym, aby ktoś pomylił informację o tym, że użyłem AI, z informacją o tym, kto ponosi odpowiedzialność za to, co właśnie przeczytał.

A to są, jak się okazuje, dwie zupełnie różne rzeczy.

I być może właśnie dlatego potrzebowaliśmy aż 200 tokenów, trzech ikonek, watermarku, metadanych, procedury, komitetu ekspertów i całego europejskiego kodeksu, żeby dojść do bardzo starego wniosku:

ostatecznie za słowo powinien odpowiadać człowiek.