Rendered at 20:48:07 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
palmotea 23 hours ago [-]
> The public posts and discussions being had about this subject are already informing AI companies on how to train their multimodal models to get around these obfuscations, most of which have already been broken. I'd argue every new font and tech demo is effectively a benchmark, daring AI firms come up with solutions to sidestep them. And they will be sidestepped, one way or another. If a human can see the information, that means there is a way the information can be parsed. "Ghost" fonts will become just another scraping obstacle with its own set of contingencies.
1. I don't like the sense of futility and powerlessness this advocates for.
2. I'm not sure it is so futile. I agree this stuff isn't encryption, which means it'll always be possible to circumvent the obfuscation, but it could raise the cost. Hopefully that can be done to the point where it's just not worth the bother.
That could happen if:
1. There are so many schemes out there the catalog of circumventions gets unwieldy.
2. Doubly so if the schemes allow generation of new obfuscated fonts per site or per page.
3. Then you're forcing the scrapers to pay a greater tax to get your text: spin up a Chrome instance to OCR a screenshot, or spend some a buck or two or LLM credits to reverse engineer the page in order to scrape it.
danudey 21 hours ago [-]
The main argument against these fonts is that you are removing accessibility for humans permanently in exchange for removing accessibility for scrapers temporarily.
It might raise the cost for some scrapers, to some degree, but it raises the cost to, effectively, infinity for anyone who needs to use a screen reader, a browser's 'reader' view, or any other assistive technology.
Things like this just remind me of EA's Spore; it released with DRM so draconian that legitimate players were getting locked out of the game during the first week, while people who pirated the game had no problem whatsoever and were playing the game without issue even before its release. The legitimate users of the thing were the only ones punished by the technology designed to stop everyone but them.
This is going to be the same thing; a website which AI will be able to read in short order but which assistive technologies will not be able to read ever.
nottorp 11 hours ago [-]
And even if you don't use a screen reader, how about search? Reader mode? Saving for later?
Oh and lynx/links...
customguy 15 hours ago [-]
So scrapers are using those people as hostages. So let's have something like a HTML meta tag to link to an accessible version of a website and hard jail for using it for scraping.
It's not a technical issue, it's a social one, people behave differently given the same set of possibilities and incentives, and we can and should target those who fuck it up for everyone.
jimmaswell 21 hours ago [-]
I also see no virtue in stopping an AI from reading something to begin with. If anything, I find it anti-social - hide your work from the machine trying to learn from it, taking nothing away from you in exchange for benefiting all of mankind?
r_lee 20 hours ago [-]
you mean benefit a private company training their AI model which they market to the public in order to get to their trillion dollar IPO?
satvikpendem 18 hours ago [-]
That's why one should advocate for open weight models.
wesleywt 11 hours ago [-]
A trillion dollar IPO buys a lot of politicians to make these illegal.
infinite_spin 20 hours ago [-]
That seems like an unfair equivalence. Lots of things you enjoy, including this very forum, benefit already wealthy private companies. That doesn't seem like it's a very good litmus test for whether something is overall a benefit to humanity.
Retric 18 hours ago [-]
> overall a benefit to humanity.
By that metric a complete ban on LLM’s might be on the table, which I don’t think is something you’re advocating for here.
infinite_spin 16 hours ago [-]
I'm advocating for a litmus test that can't be easily abused, and I'm advocating against unfair equivalences. I also am not calling "overall a benefit to humanity" a metric, it's more of a conclusion we could arrive at by an application of relevant metrics.
Retric 15 hours ago [-]
> a conclusion we could arrive at by an application of relevant metrics.
That’s dangerous ground because of how you can arbitrarily change weights of different metrics. If the goal is the “overall benefit of humanity” then that’s what’s important not metrics.
infinite_spin 14 hours ago [-]
How do you propose assessing "overall benefit of humanity" without metrics?
Retric 14 hours ago [-]
Qualitative assessment etc not just metrics.
How do you propose to asses “overall benefit to humanity” with metrics?
infinite_spin 14 hours ago [-]
Number of successful outcomes reported (e.g. with disabled students, cancer patients). Labor costs. Latency of services. Scientific discoveries found. Etc.
I'm not trying to discount the qualitative approach, but I think it's not impossibly hard to find metrics we can associate with "overall benefit to humanity" from a quantitative viewpoint.
SiempreViernes 10 hours ago [-]
> I think it's not impossibly hard to find metrics we can associate with "overall benefit to humanity" from a quantitative viewpoint.
One would hope for the slightly higher ambition of metrics that are causally connected to "overall benefit of humanity", settling for merely correlated has that unpleasant risk of the association breaking as you try to optimise.
ang_cire 8 hours ago [-]
Qualitative metrics are also metrics. You have to have a spectrum in order to weigh things against each other qualitatively, even if the spectrum is just a 2-point binary.
Retric 4 hours ago [-]
Qualitative assessment and qualitative metrics are different things.
ang_cire 4 hours ago [-]
Qualitative assessment is the overall process of using qualitative metrics to evaluate something.
infinite_spin 1 hours ago [-]
I feel like both viewpoints are valid ways of determining what you personally view as "good", and I don't want to take away from that diversity of opinion
KptMarchewa 7 hours ago [-]
There is no realistic possibility to effectively ban LLMs worldwide, which makes advocating for this useless.
SideQuark 11 hours ago [-]
So jealousy is the argument? If someone reads my work am I worse off? Why would I post it to the web? If someone indexes it so others can find it am I worse off? Should I complain they also make money helping others find it? If someone build on it in a way I did not do, never planned to do, and found really neat new tech from reading lots of works, why should I get so upset?
Seems oddly sour grapes.
neuroticnews25 10 hours ago [-]
I don't think we should hold our relations to big corporations to the same moral standards we invented for interacting with people, like "don't be jelous" or "don't be petty", because corporations won't reciprocate. They banned my account, so fuck them, simple as.
JodieBenitez 15 hours ago [-]
> taking nothing away from you
Hosting is not free.
silon42 15 hours ago [-]
That is the problem. I believe web is long overdue for a torrent like model where hosting is shared among all users (and ISPs instead of 'cloudflares').
brnt 14 hours ago [-]
Opera Unite. The idea was that hosting HTML should be as simple as browsing it.
neuroticnews25 11 hours ago [-]
Marginal cost of serving a text document is ~0.
pwdisswordfishq 11 hours ago [-]
Marginal cost of serving a text document repeatedly to hordes of reckless scrapers hiding behind residential proxies is ⋙0.
silon42 15 hours ago [-]
That is the problem. I believe web is long overdue for a torrent like model where hosting is shared among all users (and ISPs).
Alien1Being 7 hours ago [-]
Are you hopelessly naive or are you just Sam Altman ?
teekert 7 hours ago [-]
Well, as a user of LLMs I also find it annoying, the agents gather information on my behalf, they save me time and money. Are we splitting humanity? Those with info accessible to agents and those who only use their human sense directly?
This whole distinction is futile imho. And no, I'm not Sam Altman.
I mean training is one thing, you should honor people's licenses, but browsing and gathering information? Why force me to do it with my biological neural net?
tekne 5 hours ago [-]
Honestly, I'll take it a step further: what's wrong with training?
Having principles means applying them uniformly -- even to large entities or those you hate (it's fine if the principles themselves have size bounds in them, though -- versus them being implicitly glued on -- but then you need a universal justification for why that size. Which is possible and valid.)
I think that one should use the best algorithms, the best information, the best knowledge they have access to -- period. I don't like using gimped machines, I don't like making gimped machines, and I certainly don't like being sold them.
So I'm not going to turn around and say LLMs need to be gimped via arbitrary restrictions on their training data.
jimmaswell 4 hours ago [-]
I agree. As soon as I understood the gist of how modern models do what they do, I found it unreasonable to apply any stricter standards to their "learning" than we do to humans. Humans learn by reading ideas and looking at images, then go on to have ideas and draw images based on that training.
We have definitions of plagiarism and copyright infringement that apply to people based on what they put out into the world, not how they trained themselves to get there[0]. Artists literally trace art and study specific examples in detail to learn, and that's not a bad thing. Humans can also accidentally plagiarize or make things identical to past works - it's easy to think a great guitar riff just came to you when it was really based on a song you heard years ago that stuck in a part of your mind but you don't even consciously remember the influence, for example.
It would certainly be nice for AI output to provide citations if it realizes it's using a significant chunk of an idea from its training that has a clear source (or many), though this is difficult in the same way it would be difficult for me to cite where I learned about the Towers of Hanoi. The AI frequently does web searches for specific resources to get ideas from these days, which are easy for it to cite.
However, it would have a chilling effect on progress as a whole if we all started jealously guarding our ideas so close to our chest that machines couldn't read them and only a select few humans who passed some kind of gate (or even paid us) were allowed to see them. We would be nowhere close to where we are if we had always had such a mindset.
---
[0] We use proven absence of viewing certain material as a legal shield against copyright infringement e.g. clean room engineering, but this isn't strictly necessary, and possibly even discouraged with modern precedent: https://reactos.org/forum/viewtopic.php?t=21740
GlacierFox 12 hours ago [-]
Is this bait? Wtf haha
akramachamarei 12 hours ago [-]
If the comment baited you, maybe it's because you're holding onto a popular but nonsensical belief system, and you're experiencing painful cognitive dissonance.
SiempreViernes 10 hours ago [-]
Are you intentionally being rude, or is this just how HN raised you to think?
akramachamarei 3 hours ago [-]
It is not rude to point out that an insubstantial comment, which in particular falls foul of Hackernews guidelines, likely comes from a sublogical mental process. I regret my apparent failure to promote introspection or further inquiry into the issue, but I suppose the rockheadedness to which I reacted would prevent this anyway for the foreseeable future.
redsocksfan45 10 hours ago [-]
[dead]
GlacierFox 7 hours ago [-]
What a nonsensical comment. Are you experiencing painful cognitive dissonance?
pwdisswordfishq 11 hours ago [-]
> I agree this stuff isn't encryption, which means it'll always be possible to circumvent the obfuscation, but it could raise the cost. Hopefully that can be done to the point where it's just not worth the bother.
Did it work for non-cryptographic DRM? (Broadcast flag, Macrovision, deliberately miswritten floppy sectors, port dongles, physical manual challenge-response...)
palmotea 4 hours ago [-]
> Did it work for non-cryptographic DRM? (Broadcast flag, Macrovision, deliberately miswritten floppy sectors, port dongles, physical manual challenge-response...)
Yes?
Did those measures stop piracy completely? No. Did they raise the cost of piracy so there was less of it? Probably.
BrenBarn 17 hours ago [-]
I also don't like the fatalistic mindset, but I feel like we're better off trying to retaliate against the makers and operators of the bots, rather than getting into a technological arms race against the bots themselves. That is, the fight is one of policy, law, and morality, not of technology.
palmotea 14 hours ago [-]
Why not both? Better that than putting all eggs in one basket.
psd1 8 hours ago [-]
Because the effect of the technological measures is an arms race; presumably, universally so. It's a net negative. The article captures the principle in a nutshell.
palmotea 4 hours ago [-]
> Because the effect of the technological measures is an arms race; presumably, universally so. It's a net negative. The article captures the principle in a nutshell.
It's not that simple. What you say may be true, but the reason arms races happen is the alternative is surrender and domination of you or your people. You can't unilaterally choose to not participate without accepting those consequences.
evnp 1 days ago [-]
Thanks for introducing me to shieldfont.org! It's the first of these I've seen that feels designed to be more than a visual experiment, reading through their landing page is interesting. In particular, their section on accessibility seems to contradict this post's opening premise:
> Screen readers get the real words.
A screen reader reading down the page is never handed scrambled text, and our NVDA test asserts exactly that. Screen review and touch exploration are untested. ShieldFont hides shielded passages from accessibility tools by default with aria-hidden="true", because a decoy read aloud is fluent, grammatical, wrong English, and that is worse than silence.
> The real words remain sealed in the same page, and a visible notice above the block carries the control that uncovers them. It is on by default and reachable by mouse, keyboard and screen reader alike. Pressing it sets the reader’s browser to solving a compute-heavy puzzle: JavaScript and a few seconds of processing, more than most mass scrapers are willing to spend. That puts the words within reach of a screen reader, a translator and copy/paste.
I'd love to hear your thoughts on that. Also just aside, love this TUI-esque blog design and color palette (maybe a bearblog theme? still worth an upvote)
kstenerud 1 days ago [-]
When you look at their live demo (https://shieldfont.org/demo/), it says: If you use a screen reader, custom font, or translator, please uncover the text before reading.
They also actively block copying the text, telling you to "uncover" the text first. The uncover operation is VERY expensive.
Anyone using assistive technologies or trying to copy "protected" text is SOL.
Search engines will index the decoy. You'll get no traffic. Their solution is basically putting yourself in a black hole.
317070 23 hours ago [-]
> Search engines will index the decoy. You'll get no traffic. Their solution is basically putting yourself in a black hole.
It's only search engines? And who uses those anymore anyway? Other bots?
It doesn't affect organic traffic, so you are not really putting yourself into a black hole. A lot of website are driven by social media and organic traffic, so they would be just fine with this approach.
8cvor6j844qw_d6 21 hours ago [-]
Same thoughts. The black hole concerns are overstated for personal stuff when nothing is at stake. Planning to adopt one of the fonts and see how it goes.
nosioptar 6 hours ago [-]
Fortunately, the live demo doesnt work for me on Firefox mobile. I can see the text and copy/paste without "uncover".
Not sure if that's due to ublock or one of the font settings in ff.
evnp 1 days ago [-]
Isn't this the same sort of cost CloudFlare and anime catgirls are making us pay daily? Only enough to deter bots, or "a few seconds processing."
I take your point that one extra click/interaction required for screen readers only is objectively more friction, but it's only by exploring these technologies instead of dismissing them that we'll arrive at UX solutions truly work for people of all stripes (and ideally, not for bots, scrapers, and the like). I think if you're applying a tool like this you already have very different priorities than fueling Google.
kstenerud 1 days ago [-]
As soon as you need to put a "decode" button for accessibility functions to work, you're effectively posting the key along with the cipher. It's self-defeating because any tool that supports screen readers will also support scraping. It's a fool's errand.
HeatrayEnjoyer 24 hours ago [-]
Accessibility isn't optional, so what do you propose?
kstenerud 19 hours ago [-]
I propose using the internet the way it was designed: Open and readable by everything (including bots). Scraping is a legal issue. There is no technical prevention mechanism that isn't theater.
tekne 5 hours ago [-]
A better technical solution imo is more P2P hosting -- say content addressing. Something designed to deal with the scrapers -- and maybe even make them bear infrastructure costs.
mschuster91 11 hours ago [-]
> Scraping is a legal issue.
The problem is... it's a legal issue we cannot solve. America and China are large enough to not give a fuck about what everyone else wants.
csallen 1 days ago [-]
It's funny, I was just thinking that the one thing I hate most about the terminal is that its monospaced fonts and overly-long lines are a nightmare for reading. So the fact that somebody designed a blog reading experience to mimic this is just… ugh for me.
But you seem to appreciate it, and I'm sure others do too. Different strokes for different folks.
SideQuark 11 hours ago [-]
> I'd love to hear your thoughts on that.
“Claude: make the scraper mimic a screen reader.”
And just like that, in 10 seconds, their site feeds my “screen reader” the real words.
svara 14 hours ago [-]
A bit ironic that this is written in idiomatic Claudese.
Muskwalker 7 hours ago [-]
Yeah, Shieldfont's Github shows that it is being made with Claude as well, so it'll be a non-starter for those who are dogmatically anti-AI.
> Like it or not, all publicly available information will inevitably become accessible to anyone and anything that has permission to access it. This is the baseline scenario people need to plan around.
The entire point of the situation is that permission is not involved. They just do it. Meanwhile, if I do it to them, I am fined/sent to prison/executed. Until such a time that this baseline scenario of inequality is somehow remedied, there will be a motivation to stop them.
blehn 1 days ago [-]
The irony of championing accessibility using low-contrast simulated VGA text...
Narishma 1 days ago [-]
Where do you see the low contrast?
smohare 1 days ago [-]
The linked article has a fairly light gray background with white text. I can read it, but the low contrast is just tiresome.
nosioptar 6 hours ago [-]
I hate low contrast more than most.
Setting browser. display. use_document_colors or browser.display.document_color_use to 0 fixes things.
It does break some stuff, like voting buttons on hn. Im happy with the tradeoff.
22 hours ago [-]
flexagoon 10 hours ago [-]
Agree. My vision is not even that bad and yet I literally had to squint and hold my phone very close to be able to read this page.
VCFundedGenYer 1 days ago [-]
Also the notion that the sides of the pages flash as you scroll due to simulating that old Macintosh monochrome monitor dithering effect. My eyes.
boxed 14 hours ago [-]
Isn't that because you (and me!) have screens with slow response times for pixel color change though?
pibaker 7 hours ago [-]
Designing for less than ideal screens is a part of accessibility.
hn_throwaway_99 20 hours ago [-]
When I first was reading this I thought that the author was deliberately using a shitty font to make the point that obfuscated fonts are hard to read.
condour75 1 days ago [-]
Are these even meant to be used though? It seems more like performance art.
gruez 1 days ago [-]
With this kinda of stuff it's hard to tell whether the person is doing it unironically, or knows it's "performative art". A while ago there was a trend of using a tool which imperceptibly perturbs an image in a way that supposedly breaks AI training on it. Of course, artists ate it up, despite the skepticism from AI researchers. Same with people setting up their sites to be "AI scraper traps", generating gibberish content. Probably also trivial to filter out, but people do it.
pixl97 1 days ago [-]
The problem with being dumb satirically is dumb people look up to you as a thought leader.
yieldcrv 20 hours ago [-]
there was this guy that was anti-bitcoin - or specifically against most arguments from enthusiasts - and people started looking up to him to validate their feelings
and on news and podcasts he wound up correcting so many dumb arguments that he sounded pro-bitcoin and could never get to his own points
"well, no, not like that, the difficulty algorithm...."
"there are ways to use it with the power off"
"well, no, the transaction fees supplant the block reward so ..."
ffsm8 24 hours ago [-]
"caveman speak" skill, need I say more?
People aren't particularly bright. That's why the scientific method was developed to counteract our built-in tendency for... Unorthodox approaches
HlessClaudesman 14 hours ago [-]
"Go ahead, obfuscate your contribution to the repository of all human knowledge, see if that impedes our imminent invasion! Moooahahahar!!!" - Kang and Kodos
Animats 23 hours ago [-]
Agreed. There was a thing a few years ago for dazzle-painting your face to avoid face recognition. This just makes you stand out.
kazinator 17 hours ago [-]
I can tell right away that is stupid because one of the things I've used AI for was for deciphering someone's illegible handwriting, which it did amazingly.
Once it figures out for you what is written, you can't unsee it, so you know it has to be right.
Illegible fonts will only create accessibility problems for humans, even ones with normal vision, high literacy and no cognitive defects (dyslexia), while AI will blow right through the text.
If this is done in electronic documents, where the AI won't even see the glyps becaue it's reading the underlying character codes, it's even stupider.
I can't believe anyone would even try this (and then believe it is working without putting their hypotheses to the test).
simonh 11 hours ago [-]
So a rounding error of nerds, who themselves are a rounding error, will invest a lot of effort and trouble to use tech that makes their and their user's lives harder, and probably won't even work, to hide text that LLM big tech couldn't care less about anyway.
Cool.
pwdisswordfishq 11 hours ago [-]
> I don't have anything against these specific examples [...]. This post is meant to critique the idea itself. I'm not trying to put-down anyone here!
This is self-contradictory. Either you want to critique something or you do not. Make up your mind.
piker 1 days ago [-]
There could be benefits unlocked in legal documents by retaining a machine-readable version and distributing the obfuscated version with a legend at the top. We proposed one that said:
"This document contains mitigations against review by automated systems. Recipients should ensure that they have read the contents on screen or in print. Recipients with bona fide vision impairments may be entitled to unmitigated documents upon request."
In testing, obfuscating small portions of text slipped under the radar of most (then-)frontier LLMs.
We used a font that was rendered on the fly and reported faulty or fake Unicode mappings: https://tritium.legal/blog/noroboto but others have proposed and done the same with ligatures.
tbalsam 1 days ago [-]
There was a story once about a boy with a wheelchair who needed a ramp to get into school, and the school made him use the loading dock ramp used for garbage and other things at the back. The school argued that it was an appropriate accommodation.
Accessibility is not accessible if you need to go through extra steps to get it.
Cerebrally, this as a solution makes sense. But if you know anyone with a vision or other impairment, gating it behind a request is not only cruel but gets within dangerous striking distance of an ADA lawsuit, for general applications.
Maybe in the legal field or specific niche cases it's possible. But this would represent a major step backwards in the work we've done lowering barriers for a population whose only difficulty in accessing common resources is because they were born, or got sick, differently than anyone else.
piker 1 days ago [-]
My dad caught paralytic Polio at age 2 and has had limited mobility his entire life, so I'm familiar with that issue.
Our internal, hypothetical use-case was between contracting parties who were looking to avoid terms escaping into the wild. This shouldn't show up in standard ToS or similar. There are already really good legal reasons for that.
doctorpangloss 1 days ago [-]
okay, i get that as a lawyer who wants to make money, every client is "heckin cute and valid." and you can hypothesize that this thing is something that clients want: "terms escaping into the wild," whatever that means - are you saying that you think copying and pasting an agreement into an LLM makes its contents escape into the wild, by some mechanism?
Look, I understand, you don't have to explain to me the theory for how that happens, I know it already. Since I know your a smart guy, to some extent you care about that only because you imagine that clients do. But in reality, in the real world, every email you send is read by at least two people, every contract you sign has multiple parties, etc. You make some obfuscated thing or whatever, but eventually, someone has the real text of the document - it might be YOUR client, it might be the person you are negotiating with, and you rarely represent ALL the sides. You never own ALL the information and all the parties and IT systems in totality. Eventually someone will put the text into a chatbot. Or maybe they put a salient piece of the pre-final text, like some legal theory or merely a question, into the chatbot.
So I see this font stuff, or watermarking stuff, or all this provenance and control stuff, as deeply illusory. It is the worst circlejerky kind of aesthetic experience making. When you mess with anti AI fonts you are trying to compete in the same business DocuSign is in, that is, in the business of selling holistic social experiences - a whole 7000 person company whose main competition is a fucking pen - but it's not like you're doing something creative. If you care about aesthetic experiences, write a short story! Are you getting it? The itch you are scratching with this weird thing, nobody wants.
stronglikedan 1 days ago [-]
> Accessibility is not accessible if you need to go through extra steps to get it.
That seems a bit entitled to me, especially in the story you shared. They had access to the school just like everyone else. Why does it have to be in the exact same spot? Surely they could be dropped off by the loading dock as easily as others could be dropped off out front, and maybe even moreso. Should the wheelchair accessible stall be the first stall in the bathroom so they don't have to go through "extra steps" to get to it? As long as the ramp had the proper gradation to satisfy the ADA, I don't see the problem.
shnock 1 days ago [-]
> They had access to the school just like everyone else
They literally did not. They had access to the school from a different entrance than everybody else. Specifically, one intended for cargo before people.
gblargg 1 days ago [-]
They also literally cannot have the same access like everyone else, since they're in a wheelchair. Everything will be different.
fluoridation 1 days ago [-]
That's a different sense of "just like". Not "in the same manner", but "to the same extent".
binaryturtle 1 days ago [-]
Would it have been different, if they had everyone take the cargo entrance?
grim_io 1 days ago [-]
Yes. It's about discrimination and dehumanizing everyday cruelty.
psd1 8 hours ago [-]
I'm emotionally aligned with you, but the dehumanisation is the default state (e.g. the able are standing and the wheelchair user is below eye level), and all accommodations are a result of society dedicating resources to moving the needle away from that.
It may not be sufficient, and we should criticise when it isn't sufficient, but we must in practical terms also accept tradeoffs in a scarcity economy. In the worst case, some of those tradeoffs accommodate one disability over another disability.
If the dean had a swanky office and the wheelchair kid must roll past the bins, then the school made the wrong choice. But if it's a ramp at the front or textbooks, then we must think a bit harder.
22 hours ago [-]
gizmo686 1 days ago [-]
It is not sufficient to work against current AI. It needs to also work against AI that has been trained by a competent team aware of your mitigation. Or worse, a competent developer with no particular AI skills.
Otherwise, you are relying on obscurity, and will lose as soon as you become interesting enough to matter.
You will also break non-AI machine processing use cases. That isn't just accessibility, it is things like search.
piker 1 days ago [-]
Yep
aleksejs 1 days ago [-]
You will surely not have a good time enforcing the terms of a legal document that explicitly spells out that it is intentionally obfuscated from the party it intends to bind.
piker 1 days ago [-]
No, that’s not at all what is going on here.
frollogaston 18 hours ago [-]
Is the video on shieldfont.org AI-generated? It sounds like that voice again.
teekert 10 hours ago [-]
My agents act on my behalf, if you make it difficult for them I (or they) will get their info elsewhere. Will this split the internet? How long is this feasible? How long until any agent can work around any barrier, just at the expense of more energy?
rappatic 20 hours ago [-]
Ironic that an article about bad fonts would use such an ugly, garish font
hk1337 1 days ago [-]
Anti-AI fonts seems like scrambled porn on cable back in the 1980s.
jeroenhd 9 hours ago [-]
> Like it or not, all publicly available information will inevitably become accessible to anyone and anything that has permission to access it. This is the baseline scenario people need to plan around.
And this will be the end of the internet as we had it, and as we have it right now. With no way to block AI scrapers and no hope of AI companies acting ethically, shutting down the free internet is the only end result.
I'll remember these fun, bespoke attempts to stop AI from ruining the world once the internet is gone. That alone makes them worth existing for me.
woodrowbarlow 8 hours ago [-]
"become accessible to anyone that has permission to access it" is a nothing-burger sentence.
hartator 1 days ago [-]
It also mostly don’t work.
bawolff 1 days ago [-]
Yeah, it seems like most of these would probably be easier for an AI to read than a human once you give even a tiny bit of training to the AI.
there is a reason nobody uses text based captchas anymore.
theblazehen 13 hours ago [-]
Even just giving the image and "do image analysis to figure out what it says" to your LLM will get it done
spragl 13 hours ago [-]
I dont think those fonts, linked to in the article, are going to make any significant difference. But I really like the creativity behind them.
Varelion 1 days ago [-]
Is there evidence shieldfont doesn't work?
gs17 1 days ago [-]
It only works while it's rare. If it was more common, scrapers would switch to OCR or simply reverse the font so they can decode the ligatures.
Their own whitepaper brings up a bigger issue: if it works, it poisons search engines as well.
Varelion 1 days ago [-]
OCR is a more expensive option, right? I don't think these fonts need to stop ai from being trained; I think they just need to make it more expensive and difficult.
pixl97 1 days ago [-]
More expensive and not worthwhile are two different things. Also when certain implementations become popular it's much more likely someone will write a very efficient kernel for decoding said text making it much less expensive than generic OCR.
Also more expensive doesn't mean that something won't happen, only the dynamics of how it happens. For example if you put all your documents in images then some service might just sell the AI providers the text. That service may do underhanded things like bundle OCR in an app that does something else and use your phone to get the text out of these images all day.
gs17 1 days ago [-]
It's slightly more expensive, but the cost would be worth it if this became common. It's not designed to be hard to OCR, it's designed to be hard to copy out of the source code, so it doesn't require a very advanced OCR system (and all the regions that would need OCR-ing are clearly marked).
1 days ago [-]
fluoridation 1 days ago [-]
Wouldn't a font that shuffled the codepoint-to-glyph assignments be more effective?
CodesInChaos 1 days ago [-]
That wouldn't make the decoy text look plausible to the AI.
fwlr 13 hours ago [-]
>>We’re better off with the way the web is now
Well, bad news, you aren’t going to get to keep that either. Whether for high-minded reasons like “the knowledge that everything you create will be fed into the slop machine so it can later be regurgitated without attribution is having a deleterious effect on the morale of some contributors”, or for incredibly mundane reasons like “lacking the resources to either serve or block the crawlers”, I fear the web you love is dying of ai with or without anti-ai fonts.
anax32 23 hours ago [-]
Love that page styling.
tabarnacle 20 hours ago [-]
Shieldfont approaches the accessibility issue mentioned by not obfuscating text on screen readers.
NetOpWibby 15 hours ago [-]
Who's gonna be the one to bring back Flash?
yieldcrv 21 hours ago [-]
I think this is an example of just catering to the gullible solely because the market exists without pondering anything about the individuals in the market
like the "pink tax", which isn't a tax at all but just a premium on consumer gullibility as the consumer can purchase other products that do the same thing simply marketed in a different way
playing into anti AI sentiment in a useless way fits the criteria
wesleywt 11 hours ago [-]
Just give up and don't even try. This is what I am getting from this post. There are already efforts to pollute musical data that seems effective.
1 days ago [-]
dana-s 1 days ago [-]
I believe the cat and rat game is already there, for multiple places, spam, captchas and now for AI content, yes, it's objective is to make it harder for AI companies to get such data, if it wastes their time, it's a win.
rpdillon 1 days ago [-]
> it's objective is to make it harder for AI companies to get such data, if it wastes their time, it's a win.
That's only one half of the equation, though, isn't it? What if it makes it harder for legitimate users as well? It seems there's a balance to be struck.
pixl97 24 hours ago [-]
Kind of funny how you get downvoted for a rational take, but one that's not anti-ai.
You'd ask that same person how much they like captchas and I'm sure they'd think their a terrible idea and they've ran into all kinds of issues with them.
iwontberude 1 days ago [-]
[dead]
dombiscoff 15 hours ago [-]
Quite frankly I don't understand why the cat and mouse argument here is meant to serve as a shutdown for anti-AI methods. Whole industries entirely exist in a cat and mouse state (cybersecurity, anti-cheat, etc) and no one in those industries imply that any solution is or can be a permanent dunk. If anything, the very fact that theres no permanent dunk is what leads to such industries developing a competitive service based industry to begin with. Why can't a theoretical anti-AI industry develop into the same thing? These fonts just seem like the infancy steps for such.
psd1 7 hours ago [-]
Cybersecurity and KAC are not tractable to legislation because the attackers are already breaking the rules.
AI labs are broadly within the law and would like to not become criminal organisations, because money laundering is expensive.
charcircuit 12 hours ago [-]
>accessibility tools will parse. Already, you've cut out many of the humans you're trying to reach.
I figure less than 1% of the people you are trying to reach would be cut out by text that some screen readers read incorrectly.
akramachamarei 12 hours ago [-]
Well, yes, that's kinda the whole question of accessibility. The proportion of people who e.g. can't distinguish colors or use the stairs might be a rounding error in some populations. We still find it valuable to strive to serve them. Some of this has been expressed and thus calcified into regulations.
Alien1Being 7 hours ago [-]
Litigation and legislation against Big AI may be the only things that work .
We will probably have to wait Europe and China to do this since the US executive , legislative and legal systems have all been bought by Ames oligarchs .
aussieguy1234 18 hours ago [-]
It's probably trivial for AI to work around this either now or in the near future.
1. Screenshot page
2. Parse with OCR
Then train or do whatever else the font was trying to prevent...
ares623 19 hours ago [-]
Look at what they make us give.
grim_io 1 days ago [-]
DRM for your eyes == garbage idea.
waffletower 1 days ago [-]
I don't foresee anti-AI fonts being widely adopted. I see them largely as the symbolic saber rattling of intellectual property trolls.
dgellow 1 days ago [-]
Actually, pretty sure it was an art performance
KPGv2 1 days ago [-]
Every AI font proponent I've seen has been a fanfiction writer who just doesn't like AI stealing their shit to use against them.
wesleywt 11 hours ago [-]
I don't think its necessary. When everything is AI slop, people most likely will revert to trusted sources.
yehat 8 hours ago [-]
Dude, first choose a normal font for your blog.... duh
hellojomp 1 days ago [-]
We are now in a weird middle ground where we want to write things OCR algorithms have trouble transcribing which also means we write things people with accessibility issues have trouble seeing. No child left behind?
mister_mort 1 days ago [-]
It's like the old tale about the national park bin with the smartest bear / dumbest tourist crossover, except we're now comparing capabilities of the smartest AI with disabled humans.
Xirdus 1 days ago [-]
It already was a major issue in 2010 - home desktop-grade OCRs could easily beat an average grandma on reading heavily garbled text.
no-name-here 1 days ago [-]
> in 2010 - home desktop-grade OCRs could easily beat an average grandma
I’ve heard Tesseract OCR recommended, but even in 2026, no matter how I scan receipts or documents, the OCR output seems far worse than human reading?
strangecasts 1 days ago [-]
I think the difficulty was specifically with CAPTCHA challenges, which had to be quick to generate but still legible - OCR on physical documents has to be robust to a different set of problems
That said, what kind of errors are you getting? I think the main practical difference between Tesseract's LSTM-based OCR and newer VLM-based OCR like PaddleOCR [1] is (hopefully) getting to skip making heuristics for the layout of the document, but errors possibly compounding over multiple tokens - are you getting individual illegible words or having the documents smooshed together because the OCR can't parse the layout?
>> no matter how I scan receipts or documents, the OCR output seems far worse than human
> what kind of errors are you getting?
Just noticeably worse character recognition, particularly where the document is faded, water-stained, or the paper document (not the scan) was low resolution to begin with, as compared to 'normal' human recognition.
I was largely using Tesseract in conjunction with the self-hosted Paperless-NGX, and I wanted to stay free/local without yet investing in AI-focused hardware. But you're right that AI will continue to advance, including in smaller local models, and it looks like Paperless 3.0 released recently (after my testing earlier in 2026), including with AI functionality.
> or having the documents smooshed together because the OCR can't parse the layout
I haven't even been worrying about that yet - I'm just at the point of trying to get the OCR characters right. :-)
Xirdus 1 days ago [-]
Well, I used specialized software specifically for solving CAPTCHA on select few file sharing websites. My later experience with general purpose OCR software was just as awful as yours. But it's a product problem, not technology problem - you really don't need an advanced AI for it, you just need a good implementation of traditional shape-matching OCR that doesn't seem to exist anywhere for some reason.
julianfurchert 13 hours ago [-]
[dead]
unethical_ban 1 days ago [-]
Is it part of the joke that the site is intentionally over-pixelated while the author critiques readability? (edit: I don't mind esoteric design and I play old games. I found it funny to see a blog with aesthetics that are not optimized for long reading to complain about the readability of fonts)
I don't think any anti-AI font design is more than design-as-art statements against AI. If there is evidence of them being used exclusively for business and without accessible fonts aside them, I'm willing to be wrong.
I'm aware, and my comment isn't meant to be a challenge to anyone's sacred honor. I was observing a juxtaposition.
avazhi 1 days ago [-]
Not everything that annoys or inconveniences you is harmful, as if this needs to be said to an adult.
bradthebeaverfa 21 hours ago [-]
This author greatly overestimates how much I care about accessibility.
If I have to block a blind person from reading my blog to block an AI from training on it, I'll make that trade every time.
Ohentis 21 hours ago [-]
Well of the fonts listed, 2 out of the 3 would also prevent any human from wanting to read it and the third becomes an ineffective counter measure if it's widely used.
dombiscoff 14 hours ago [-]
I think the line of reasoning that a method should provide as much accessibility as possible to not alienate real humans is a good one. You use 'have to' as if there aren't better solutions that could be invented, but such will only occur if we challenge the imperfect solutions we have now.
phoghed 18 hours ago [-]
I don’t understand how you think a font would block a bot from reading your blog in the first place?
It’s not even going to render the damn website. If it decides to and detects your retarded font it could just change the font trivially. You’d have to fundamentally fuck up the html text content for it to work at all, and then all you’ll do is inconvenience real people that you’d be extremely lucky to have attracted to your blog in the first place.
hyperadvanced 18 hours ago [-]
For real. It’s such a weak cop-out of an argument to lead with that I clicked out, never to read this blog again.
yipinwong 1 days ago [-]
Not trynna to be funny.
The author's post is anti-human, thus useless and harmful (medically).
I can't read this bad font, sizing, spacing, etc.
The main offender is the color choice, and fonts that are just god aweful to read.
dwrodri 1 days ago [-]
Accessibility is important, and I find that "Reader Mode" in most browsers is quite good. Everyone should have access to the tools to consume content. Did your browser not provide that functionality?
yipinwong 1 days ago [-]
basically make it look hard to read for the majority for the sake of 1%?
let those 1% use a diff tool view instead of 99% of us having to suffer.
That website no-way is accessible for my 30 year old eyes.
dombiscoff 14 hours ago [-]
I think he just enjoys terminal styling, man.
jotux 1 days ago [-]
Found it awful to look at, tried to zoom out and everything on the page got larger. Seriously gross accessibility.
kokanee 1 days ago [-]
I'm a bit frustrated by what seems to be a widespread strong negative reaction to anti-AI fonts. The accessibility problem is real, but I feel like that's a reason to push the investigation deeper for solutions to that problem, not a reason to abandon the effort entirely. The largest intellectual property infringement in the history of the universe is actively unfolding, and it's resulting in an existentially threatening transfer of wealth and power. That's a problem worth exploring every solution for, and solving it may entail some serious sacrifices.
Uberzi 1 days ago [-]
It's simply that a font won't solve anything. That concept is worst than security by obscurity, as it causes more problem and add more constraints than what it solves... for a very limited time until AI bots are adjusted to decode those fonts properly.
waterTanuki 20 hours ago [-]
how can you make such a bold claim so early with 0 evidence? A successful (and much more realistic outcome) could be to make a font that is financially infeasible for LLMs to parse but easy for humans.
psd1 7 hours ago [-]
Broadly, by reasoning from general principles, and applying a confidence level.
I suppose Uberzi has a confidence level in their claim, but it's tedious to hedge every damn thing.
SideQuark 11 hours ago [-]
Yeah, that worked really well for captchas. Now even simple models far surpass humans at recognition.
So there’s plenty of evidence.
Now provide evidence it’s technically possible to make any font “ that is financially infeasible for LLMs to parse but easy for humans.”
Every half assed means of trying to confuse an AI is just a small bit of learning away from making the AI better than you.
Worse when you have people with disabilities, which I seem to be this week, you just make doing things a pain in the ass.
What I don't get is people like you think there is a solution to this. There is not. The harder you try you either exclude more actual humans or you align the AI closer to how people actually see.
1. I don't like the sense of futility and powerlessness this advocates for.
2. I'm not sure it is so futile. I agree this stuff isn't encryption, which means it'll always be possible to circumvent the obfuscation, but it could raise the cost. Hopefully that can be done to the point where it's just not worth the bother.
That could happen if:
1. There are so many schemes out there the catalog of circumventions gets unwieldy.
2. Doubly so if the schemes allow generation of new obfuscated fonts per site or per page.
3. Then you're forcing the scrapers to pay a greater tax to get your text: spin up a Chrome instance to OCR a screenshot, or spend some a buck or two or LLM credits to reverse engineer the page in order to scrape it.
It might raise the cost for some scrapers, to some degree, but it raises the cost to, effectively, infinity for anyone who needs to use a screen reader, a browser's 'reader' view, or any other assistive technology.
Things like this just remind me of EA's Spore; it released with DRM so draconian that legitimate players were getting locked out of the game during the first week, while people who pirated the game had no problem whatsoever and were playing the game without issue even before its release. The legitimate users of the thing were the only ones punished by the technology designed to stop everyone but them.
This is going to be the same thing; a website which AI will be able to read in short order but which assistive technologies will not be able to read ever.
Oh and lynx/links...
It's not a technical issue, it's a social one, people behave differently given the same set of possibilities and incentives, and we can and should target those who fuck it up for everyone.
By that metric a complete ban on LLM’s might be on the table, which I don’t think is something you’re advocating for here.
That’s dangerous ground because of how you can arbitrarily change weights of different metrics. If the goal is the “overall benefit of humanity” then that’s what’s important not metrics.
How do you propose to asses “overall benefit to humanity” with metrics?
I'm not trying to discount the qualitative approach, but I think it's not impossibly hard to find metrics we can associate with "overall benefit to humanity" from a quantitative viewpoint.
One would hope for the slightly higher ambition of metrics that are causally connected to "overall benefit of humanity", settling for merely correlated has that unpleasant risk of the association breaking as you try to optimise.
Seems oddly sour grapes.
Hosting is not free.
This whole distinction is futile imho. And no, I'm not Sam Altman.
I mean training is one thing, you should honor people's licenses, but browsing and gathering information? Why force me to do it with my biological neural net?
Having principles means applying them uniformly -- even to large entities or those you hate (it's fine if the principles themselves have size bounds in them, though -- versus them being implicitly glued on -- but then you need a universal justification for why that size. Which is possible and valid.)
I think that one should use the best algorithms, the best information, the best knowledge they have access to -- period. I don't like using gimped machines, I don't like making gimped machines, and I certainly don't like being sold them.
So I'm not going to turn around and say LLMs need to be gimped via arbitrary restrictions on their training data.
We have definitions of plagiarism and copyright infringement that apply to people based on what they put out into the world, not how they trained themselves to get there[0]. Artists literally trace art and study specific examples in detail to learn, and that's not a bad thing. Humans can also accidentally plagiarize or make things identical to past works - it's easy to think a great guitar riff just came to you when it was really based on a song you heard years ago that stuck in a part of your mind but you don't even consciously remember the influence, for example.
It would certainly be nice for AI output to provide citations if it realizes it's using a significant chunk of an idea from its training that has a clear source (or many), though this is difficult in the same way it would be difficult for me to cite where I learned about the Towers of Hanoi. The AI frequently does web searches for specific resources to get ideas from these days, which are easy for it to cite.
However, it would have a chilling effect on progress as a whole if we all started jealously guarding our ideas so close to our chest that machines couldn't read them and only a select few humans who passed some kind of gate (or even paid us) were allowed to see them. We would be nowhere close to where we are if we had always had such a mindset.
---
[0] We use proven absence of viewing certain material as a legal shield against copyright infringement e.g. clean room engineering, but this isn't strictly necessary, and possibly even discouraged with modern precedent: https://reactos.org/forum/viewtopic.php?t=21740
Did it work for non-cryptographic DRM? (Broadcast flag, Macrovision, deliberately miswritten floppy sectors, port dongles, physical manual challenge-response...)
Yes?
Did those measures stop piracy completely? No. Did they raise the cost of piracy so there was less of it? Probably.
It's not that simple. What you say may be true, but the reason arms races happen is the alternative is surrender and domination of you or your people. You can't unilaterally choose to not participate without accepting those consequences.
> Screen readers get the real words. A screen reader reading down the page is never handed scrambled text, and our NVDA test asserts exactly that. Screen review and touch exploration are untested. ShieldFont hides shielded passages from accessibility tools by default with aria-hidden="true", because a decoy read aloud is fluent, grammatical, wrong English, and that is worse than silence.
> The real words remain sealed in the same page, and a visible notice above the block carries the control that uncovers them. It is on by default and reachable by mouse, keyboard and screen reader alike. Pressing it sets the reader’s browser to solving a compute-heavy puzzle: JavaScript and a few seconds of processing, more than most mass scrapers are willing to spend. That puts the words within reach of a screen reader, a translator and copy/paste.
I'd love to hear your thoughts on that. Also just aside, love this TUI-esque blog design and color palette (maybe a bearblog theme? still worth an upvote)
They also actively block copying the text, telling you to "uncover" the text first. The uncover operation is VERY expensive.
Anyone using assistive technologies or trying to copy "protected" text is SOL.
Search engines will index the decoy. You'll get no traffic. Their solution is basically putting yourself in a black hole.
It's only search engines? And who uses those anymore anyway? Other bots?
It doesn't affect organic traffic, so you are not really putting yourself into a black hole. A lot of website are driven by social media and organic traffic, so they would be just fine with this approach.
Not sure if that's due to ublock or one of the font settings in ff.
I take your point that one extra click/interaction required for screen readers only is objectively more friction, but it's only by exploring these technologies instead of dismissing them that we'll arrive at UX solutions truly work for people of all stripes (and ideally, not for bots, scrapers, and the like). I think if you're applying a tool like this you already have very different priorities than fueling Google.
The problem is... it's a legal issue we cannot solve. America and China are large enough to not give a fuck about what everyone else wants.
But you seem to appreciate it, and I'm sure others do too. Different strokes for different folks.
“Claude: make the scraper mimic a screen reader.”
And just like that, in 10 seconds, their site feeds my “screen reader” the real words.
Their stance appears to be "[t]his is not an anti-AI font. AI is in our lives. We are actually a pro-consent font." (https://github.com/isaqueseneda/shieldfont/issues/2) Might be a bit more niche market.
The entire point of the situation is that permission is not involved. They just do it. Meanwhile, if I do it to them, I am fined/sent to prison/executed. Until such a time that this baseline scenario of inequality is somehow remedied, there will be a motivation to stop them.
Setting browser. display. use_document_colors or browser.display.document_color_use to 0 fixes things.
It does break some stuff, like voting buttons on hn. Im happy with the tradeoff.
and on news and podcasts he wound up correcting so many dumb arguments that he sounded pro-bitcoin and could never get to his own points
"well, no, not like that, the difficulty algorithm...."
"there are ways to use it with the power off"
"well, no, the transaction fees supplant the block reward so ..."
People aren't particularly bright. That's why the scientific method was developed to counteract our built-in tendency for... Unorthodox approaches
Illegible fonts will only create accessibility problems for humans, even ones with normal vision, high literacy and no cognitive defects (dyslexia), while AI will blow right through the text.
If this is done in electronic documents, where the AI won't even see the glyps becaue it's reading the underlying character codes, it's even stupider.
I can't believe anyone would even try this (and then believe it is working without putting their hypotheses to the test).
Cool.
This is self-contradictory. Either you want to critique something or you do not. Make up your mind.
"This document contains mitigations against review by automated systems. Recipients should ensure that they have read the contents on screen or in print. Recipients with bona fide vision impairments may be entitled to unmitigated documents upon request."
In testing, obfuscating small portions of text slipped under the radar of most (then-)frontier LLMs.
We used a font that was rendered on the fly and reported faulty or fake Unicode mappings: https://tritium.legal/blog/noroboto but others have proposed and done the same with ligatures.
Accessibility is not accessible if you need to go through extra steps to get it.
Cerebrally, this as a solution makes sense. But if you know anyone with a vision or other impairment, gating it behind a request is not only cruel but gets within dangerous striking distance of an ADA lawsuit, for general applications.
Maybe in the legal field or specific niche cases it's possible. But this would represent a major step backwards in the work we've done lowering barriers for a population whose only difficulty in accessing common resources is because they were born, or got sick, differently than anyone else.
Our internal, hypothetical use-case was between contracting parties who were looking to avoid terms escaping into the wild. This shouldn't show up in standard ToS or similar. There are already really good legal reasons for that.
Look, I understand, you don't have to explain to me the theory for how that happens, I know it already. Since I know your a smart guy, to some extent you care about that only because you imagine that clients do. But in reality, in the real world, every email you send is read by at least two people, every contract you sign has multiple parties, etc. You make some obfuscated thing or whatever, but eventually, someone has the real text of the document - it might be YOUR client, it might be the person you are negotiating with, and you rarely represent ALL the sides. You never own ALL the information and all the parties and IT systems in totality. Eventually someone will put the text into a chatbot. Or maybe they put a salient piece of the pre-final text, like some legal theory or merely a question, into the chatbot.
So I see this font stuff, or watermarking stuff, or all this provenance and control stuff, as deeply illusory. It is the worst circlejerky kind of aesthetic experience making. When you mess with anti AI fonts you are trying to compete in the same business DocuSign is in, that is, in the business of selling holistic social experiences - a whole 7000 person company whose main competition is a fucking pen - but it's not like you're doing something creative. If you care about aesthetic experiences, write a short story! Are you getting it? The itch you are scratching with this weird thing, nobody wants.
That seems a bit entitled to me, especially in the story you shared. They had access to the school just like everyone else. Why does it have to be in the exact same spot? Surely they could be dropped off by the loading dock as easily as others could be dropped off out front, and maybe even moreso. Should the wheelchair accessible stall be the first stall in the bathroom so they don't have to go through "extra steps" to get to it? As long as the ramp had the proper gradation to satisfy the ADA, I don't see the problem.
They literally did not. They had access to the school from a different entrance than everybody else. Specifically, one intended for cargo before people.
It may not be sufficient, and we should criticise when it isn't sufficient, but we must in practical terms also accept tradeoffs in a scarcity economy. In the worst case, some of those tradeoffs accommodate one disability over another disability.
If the dean had a swanky office and the wheelchair kid must roll past the bins, then the school made the wrong choice. But if it's a ramp at the front or textbooks, then we must think a bit harder.
Otherwise, you are relying on obscurity, and will lose as soon as you become interesting enough to matter.
You will also break non-AI machine processing use cases. That isn't just accessibility, it is things like search.
And this will be the end of the internet as we had it, and as we have it right now. With no way to block AI scrapers and no hope of AI companies acting ethically, shutting down the free internet is the only end result.
I'll remember these fun, bespoke attempts to stop AI from ruining the world once the internet is gone. That alone makes them worth existing for me.
there is a reason nobody uses text based captchas anymore.
Their own whitepaper brings up a bigger issue: if it works, it poisons search engines as well.
Also more expensive doesn't mean that something won't happen, only the dynamics of how it happens. For example if you put all your documents in images then some service might just sell the AI providers the text. That service may do underhanded things like bundle OCR in an app that does something else and use your phone to get the text out of these images all day.
Well, bad news, you aren’t going to get to keep that either. Whether for high-minded reasons like “the knowledge that everything you create will be fed into the slop machine so it can later be regurgitated without attribution is having a deleterious effect on the morale of some contributors”, or for incredibly mundane reasons like “lacking the resources to either serve or block the crawlers”, I fear the web you love is dying of ai with or without anti-ai fonts.
like the "pink tax", which isn't a tax at all but just a premium on consumer gullibility as the consumer can purchase other products that do the same thing simply marketed in a different way
playing into anti AI sentiment in a useless way fits the criteria
That's only one half of the equation, though, isn't it? What if it makes it harder for legitimate users as well? It seems there's a balance to be struck.
You'd ask that same person how much they like captchas and I'm sure they'd think their a terrible idea and they've ran into all kinds of issues with them.
AI labs are broadly within the law and would like to not become criminal organisations, because money laundering is expensive.
I figure less than 1% of the people you are trying to reach would be cut out by text that some screen readers read incorrectly.
We will probably have to wait Europe and China to do this since the US executive , legislative and legal systems have all been bought by Ames oligarchs .
1. Screenshot page
2. Parse with OCR
Then train or do whatever else the font was trying to prevent...
I’ve heard Tesseract OCR recommended, but even in 2026, no matter how I scan receipts or documents, the OCR output seems far worse than human reading?
That said, what kind of errors are you getting? I think the main practical difference between Tesseract's LSTM-based OCR and newer VLM-based OCR like PaddleOCR [1] is (hopefully) getting to skip making heuristics for the layout of the document, but errors possibly compounding over multiple tokens - are you getting individual illegible words or having the documents smooshed together because the OCR can't parse the layout?
[1] https://github.com/PaddlePaddle/PaddleOCR
>> no matter how I scan receipts or documents, the OCR output seems far worse than human
> what kind of errors are you getting?
Just noticeably worse character recognition, particularly where the document is faded, water-stained, or the paper document (not the scan) was low resolution to begin with, as compared to 'normal' human recognition.
I was largely using Tesseract in conjunction with the self-hosted Paperless-NGX, and I wanted to stay free/local without yet investing in AI-focused hardware. But you're right that AI will continue to advance, including in smaller local models, and it looks like Paperless 3.0 released recently (after my testing earlier in 2026), including with AI functionality.
> or having the documents smooshed together because the OCR can't parse the layout
I haven't even been worrying about that yet - I'm just at the point of trying to get the OCR characters right. :-)
I don't think any anti-AI font design is more than design-as-art statements against AI. If there is evidence of them being used exclusively for business and without accessible fonts aside them, I'm willing to be wrong.
If I have to block a blind person from reading my blog to block an AI from training on it, I'll make that trade every time.
It’s not even going to render the damn website. If it decides to and detects your retarded font it could just change the font trivially. You’d have to fundamentally fuck up the html text content for it to work at all, and then all you’ll do is inconvenience real people that you’d be extremely lucky to have attracted to your blog in the first place.
I can't read this bad font, sizing, spacing, etc. The main offender is the color choice, and fonts that are just god aweful to read.
That website no-way is accessible for my 30 year old eyes.
I suppose Uberzi has a confidence level in their claim, but it's tedious to hedge every damn thing.
So there’s plenty of evidence.
Now provide evidence it’s technically possible to make any font “ that is financially infeasible for LLMs to parse but easy for humans.”
Every half assed means of trying to confuse an AI is just a small bit of learning away from making the AI better than you.
Worse when you have people with disabilities, which I seem to be this week, you just make doing things a pain in the ass.
What I don't get is people like you think there is a solution to this. There is not. The harder you try you either exclude more actual humans or you align the AI closer to how people actually see.