Reading view

The web’s newest weapon against AI scrapers is a font

AI companies' penchant for scraping through large swaths of the public web in search of valuable training data has already led to lawsuits and technical fixes aimed at stopping the practice. Now, a pair of designers is hoping to stymie these scrapers with a new font designed to offer people a perfectly readable webpage while serving scrapers a subtly edited, nonsensical version in the underlying HTML.

ShieldFont, as designers Isaque Seneda and Gabriel Abrucio write in a recent white paper, was made to offer web publishers "a practical opt-out from unauthorized AI training and [to] disrupt what is collected when that choice is ignored."

When is a horse a potato?

The font is based around ligatures, a long-standing feature of many fonts that is usually used to replace certain letter pairs with a more readable version when they're smushed up next to each other. With ShieldFont, though, those ligatures are instead used to replace entire words with others in an attempt to destroy the text's value to scrapers. This substitution only happens when the font engine draws the page onscreen, meaning scrapers that simply download plaintext source code get an altered version that end users never see.

Read full article

Comments

© Getty Images

  •  

Booksellers suspect AI firms are buying and then destroying rare books

If you can truly appreciate an old book—and maybe even marvel at how its fragile, yellowing pages contain some of the earliest ways that people tried to make sense of the world around them—then headlines about tech companies that are destroying books to train AI likely torture a tender part of your soul.

It’s indeed depressing to imagine piles of book spines waiting to be fed into wood chippers while torn-out pages are cropped, scanned, and trashed. But that’s the cheapest and easiest way to scan books as fast as possible, and AI companies are in a race to advance their models by training on the kind of engaging, high-quality long-form texts that can only be found in books. So book lovers fear it’s likely that the practice is happening on a grander scale than is currently being reported and that some physical copies of books will be lost forever.

What makes this destruction extra painful, though, is that it doesn’t have to be this way.

Read full article

Comments

© hexvivo | iStock / Getty Images Plus

  •  

With new open models, Meta pitches another reboot of its struggling AI strategy

Meta has announced its intention to focus on open-weight large language models. Additionally, the company announced the release of an open model called Muse Glimmer and a promise to open the weights for Muse Spark 1.2, its more powerful model, in the next few weeks.

Alongside these announcements, Meta CEO Mark Zuckerberg published a more than 6,000-word essay outlining the company's philosophy about AI systems and governance moving forward. The essay aims to differentiate Meta from companies like OpenAI and Anthropic, which develop proprietary models and have lobbied the US government for help competing against large-scale distillation—which involves using an existing model to train a new one—or open-weight models by Chinese labs.

Read full article

Comments

© Bloomberg / Contributor | Bloomberg

  •  
❌