cross-posted from: https://lemmy.world/post/49778353

Also /c/modabuse. Also /c/yepowertrippinbastards. Also lemmy.ml/c/worldnews, comrade, ACAB, and 552 individual accounts across 67 instances, about half of them on lemmy.world.

None of this is in the source code. It’s downloaded at runtime from a file nobody has ever looked at.

If you’re just tuning in

Tesseract is a third-party web frontend for Lemmy, maintained by asimons04 and licensed AGPL-3.0. Admins deploy it on their own servers alongside or instead of lemmy-ui, and there are public instances of it people use to browse Lemmy generally. If you’ve used a Lemmy site that didn’t look like stock Lemmy, there’s a fair chance it was this.

Last week db0 posted a PSA: Tesseract contains a blacklist of instance domains compiled directly into the application. 32 of them. Admins can’t see it, can’t configure it, and aren’t told it’s there. Connect to a listed instance and the app tells you it’s “incompatible,” which is not true.

I went through the code to see how that was implemented. The hardcoded list turns out to be the small half of the system.

There’s a second filter policy fetched over HTTP every time the app loads. It isn’t in the git repository. It’s unauthenticated and world-readable, so anyone can pull it. Right now it carries 552 user accounts, 2,275 username patterns, 54 instances, 97 communities, 289 keyword patterns and 351 domains, with every category set to hide matches rather than flag them. Not collapsed behind a click. Simply absent, with no indication anything was removed.

Verify all of it in ten seconds

curl -s https://tesseract.dubvee.org/tesseract/api/system/policy \
  | base64 -d | gunzip > policy.json

That’s the live policy, base64-wrapped gzip, 111KB of JSON when it unpacks. There’s a stale fallback copy at /data/policy.dat as well.

It filters criticism of moderators

  • lemmy.sdf.org/c/modabuse — listed
  • lemmy.dbzer0.com/c/yepowertrippinbastards — listed
  • lemmy.dbzer0.com/c/YPTBcirclejerk — listed
  • community regex power ?tripping?
  • keyword censoring me

Call the rest of it whatever you like. This part is not spam defence.

It filters words

The 32 community name patterns include Communis(t|m), Conservativ(e|es|ism), Leftis(t|m), Libertarian(ism)?, ^Green Part(y|ies), Zionis(t|m), (Police|Cops), guillotine and billionaire.

Keywords include comrade, ACAB, neoliberal, proletaria(n|t) and death to.

Filtered communities on instances that aren’t blocked: lemmy.ml/c/worldnews, lemmy.today/c/news, lemmy.ca/c/politicalnewscanada, lemmy.ca/c/usa, infosec.pub/c/strategic_unions.

The 552 users aren’t bots

67 instances. 272 on lemmy.world alone, 40 on sh.itjust.works, 19 on lemmy.ca, and 28 instances contributing exactly one person each.

355 of the 552 usernames are plain alphabetic, twelve characters or under, median length eight. Only 36 look like spam registrations. A bot list looks like the opposite of that.

Seven of them aren’t even Lemmy. There are Mastodon and Friendica accounts in there: people who have never used Lemmy, hidden by a Lemmy frontend, with no possible way of finding out.

I have the list and I’m not posting it. Most of these are ordinary people who got pattern-matched, and 552 names on this comm is a harassment target inside an hour. Run the command above and grep for yourself.

And it lies about it

When the instance block fires you get: “Incompatible Instance. Not Supported. $instance is not compatible with Tesseract.”

Nothing is incompatible. It’s a policy decision dressed as an API error, and it’s what had db0 chasing a version mismatch that never existed.

For the hidden users, communities and keywords, you get no message at all.

Admins can’t switch it off

Tesseract has env vars for PUBLIC_DOMAIN_BLACKLIST, PUBLIC_FAKE_NEWS_BLACKLIST and the shortener lists. There is none for either blocklist. enableToxicMode bypasses the other filters and explicitly not this one.

Self-host it and you cannot disable this, nothing in your config admits it exists, and the contents can change without you pulling a commit.

Before someone says it

A lot of that domain list is real spam defence. It filters conservatism as well as communism. “It targets the left” doesn’t survive the data and I’m not going to pretend it does.

The problem is that spam filtering and political editorial got welded into one undocumented, remotely-updatable blob, shipped hidden, to admins who’ve never read it and users who don’t know it’s there. The spam work is what makes the rest unauditable: “it’s a spam list” answers every individual question and none of the whole.

And /c/modabuse is not spam.

Asks

  1. Publish the runtime policy in the repo, or kill the endpoint.
  2. Stop reporting a policy block as a technical incompatibility.
  3. Tell users when something’s been hidden. One line.
  4. Give operators an off switch, like every other blacklist in the codebase has.

It’s AGPL-3.0 and db0 already forked it. That’s the licence working as designed. But forking isn’t disclosure, and the admins who need this are precisely the ones with no reason to go looking.

If you run Tesseract, you are relaying a 111KB moderation policy you have never read, under your instance’s name, to users who don’t know it exists.

Full contents of every list, unedited, in the comments.

  • subignition@fedia.io
    link
    fedilink
    arrow-up
    9
    arrow-down
    4
    ·
    20 hours ago

    Hi, not the person you replied to, and (edit: not*) agreeing with any of this blocklist shit, AND my ire is firmly not shifted away from Tesseract, but just wanted to break down the characteristics of the OP that suggest (to me) it was composed mostly with an LLM and fine tuned after. (Though to be clear, that’s a pretty unimportant characteristic of the post considering the actual content)

    1. The division of the text into small “sections” of five or fewer short paragraphs separated by headers is probably the most obvious hallmark.

    2. Kind of difficult to put into words, but abrupt / awkward / overly poetic metaphor and… “overconfident”? “sensationalized”? phrasing both give me an LLM vibe. examples:

    It’s a policy decision dressed as an API error, and it’s what had db0 chasing a version mismatch that never existed.

    I went through the code to see how that was implemented. The hardcoded list turns out to be the small half of the system.

    The spam work is what makes the rest unauditable: “it’s a spam list” answers every individual question and none of the whole.

    It’s downloaded at runtime from a file nobody has ever looked at.

    spam filtering and political editorial got welded into one undocumented, remotely-updatable blob, shipped hidden

    That’s the licence working as designed. But forking isn’t disclosure, and the admins who need this are precisely the ones with no reason to go looking.

    1. Overuse of the rule of three. This one is probably the most tenuous, because genuine human writing does make frequent use of the rule of three. But when LLMs do it, it’s usually at a high enough density where it starts sounding like it’s trying too hard, or it’s clickbaity.

    Admins can’t see it, can’t configure it, and aren’t told it’s there.

    That’s the live policy, base64-wrapped gzip, 111KB of JSON when it unpacks.

    There are Mastodon and Friendica accounts in there: people who have never used Lemmy, hidden by a Lemmy frontend, with no possible way of finding out.

    355 of the 552 usernames are plain alphabetic, twelve characters or under, median length eight.

    For the hidden users, communities and keywords, you get no message at all.

    Self-host it and you cannot disable this, nothing in your config admits it exists, and the contents can change without you pulling a commit.

    If you run Tesseract, you are relaying a 111KB moderation policy you have never read, under your instance’s name, to users who don’t know it exists.

    1. Tonally inconsistent with the intended audience. I realize Lemmy/Fediverse users are more technically savvy than usual, but even then… consider the claim “I went through the code to see how that was implemented.” early in the post. A lot of plausible technical jargon is used, and there are plenty examples of the actual filters and filter patterns themselves, but the post makes claims about what is happening without really being specific at all about how.

    Could be explained by me just not being familiar enough with the technical side, but saying stuff like “There’s a second filter policy fetched over HTTP every time the app loads” without breaking it down any further seems a bit suspicious to me? It doesn’t actually call out any function, file, or… i dunno, a logical starting point, for a competent reader to dig into the source themselves. Reads like it’s written for either an expert who doesn’t need any details, or for an audience with zero skepticism whatsoever.

    1. And least importantly, because my focus is on the presentation of the post moreso than the accuracy, IF rimu’s read of the code is correctly calling out a factual error in the OP (and let me be clear that I’m not enough of a programmer to evaluate whether it is), that is pretty strong evidence of an LLM getting something wrong.

    Again… this wall of text isn’t meant as a refutation of anything you’re saying, or a reason to dismiss the contents of the OP uncritically. Seems pretty clear after reading through db0’s posts about it that this is seriously fucked up. I hope I’ve been clear enough about that.

    • Zawaj_Al_Qasirat@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      6
      ·
      19 hours ago

      Could be explained by me just not being familiar enough with the technical side, but saying stuff like “There’s a second filter policy fetched over HTTP every time the app loads” without breaking it down any further seems a bit suspicious to me? It doesn’t actually call out any function, file, or… i dunno, a logical starting point, for a competent reader to dig into the source themselves. Reads like it’s written for either an expert who doesn’t need any details, or for an audience with zero skepticism whatsoever.

      It gives you the file directly

      curl -s https://tesseract.dubvee.org/tesseract/api/system/policy
      | base64 -d | gunzip > policy.json

        • Zawaj_Al_Qasirat@sh.itjust.works
          link
          fedilink
          English
          arrow-up
          4
          ·
          17 hours ago

          ./Dockerfile:14:COPY --chown=node:node static/data/policy.dat /app/system/policy/policy.dat
          ./src/lib/policies/system.ts:105: res = await fetch(‘/data/policy.dat’)
          ./src/server/system.ts:27: let file = await open(${policyDir}/policy.dat)
          ./src/server/system.ts:71: await writeFile(${policyDir}/policy.dat, policy)

          • ripcord@lemmy.world
            link
            fedilink
            English
            arrow-up
            2
            ·
            14 hours ago

            But none of those are the remotely fetched file.

            There seems to ALSO be a statically included policy file with a bunch of stuff. That he updates periodically.

            • subignition@fedia.io
              link
              fedilink
              arrow-up
              1
              ·
              10 hours ago

              Give them a little credit, they were only five lines away from a bullseye in their second example. That’s much closer to accurate than their overall comprehension of my post was

          • subignition@fedia.io
            link
            fedilink
            arrow-up
            1
            arrow-down
            1
            ·
            16 hours ago

            Thanks, but I’m not sure what value you think you’re adding by actually going and finding that - if you had understood my post, you would know that my critique of the lack of specifics in the OP was a reason to doubt it was written by a human, not anything to do with whether or not the details of the post were credible or not.

            edit: the point is, including the filter enough was enough to make the post contents credible, but if it had specifically followed up with something like “and this is retrieved at src/lib/policies/system.ts:100” it would have seemed like it was written by a human and not an LLM.