Engines
Keeping engine definitions current, and engine bugs found while testing. Part of Experiments and decisions. Newest notes go at the top of each section.
Keeping platform and engine maintenance current (2026-09-29)
- Firefox minimum: AMO lint warned that Firefox for Android 140 predates support for the built-in data-collection manifest key. Raised
gecko.strict_min_versionto 142, the first Android version with that key; deliberately did not addgecko_android, which would advertise untested phone support. The generated Firefox manifest now declares 142.0 andweb-ext lintreports zero warnings. - Action runtimes: CI reported the pinned GitHub Actions running on the removed Node 20 runtime. Replaced the checkout and setup-node pins across CI, docs, releases, and the SERPINFO watch with immutable SHAs for releases whose metadata specifies Node 24.
- Engine selector ownership: moved the Google, DuckDuckGo, and Brave cleanup selectors into each
EngineDef. This removes the parallel engine-id registry and keeps a selector next to the result layout it describes. SERPINFO remains a weekly review signal, not a selector generator: its model differs from Anubis's structural detection, boundaries, and cleanup rules.
Keeping engine definitions in sync (2026-09-29)
- Starting point: the weekly workflow already watched eight uBlacklist
serpinfo/*.ymlfiles, but its file-to-engine mapping lived only inwatch-engines.mjs. It could silently fall behind when an engine was added, and fetching the 50 most recent engine issues could miss a previously reported commit. - Upstream check: GitHub's Contents API listed the current
serpinfo/files, includinggoogle.yml,duckduckgo.yml,bing.yml,brave.yml,startpage.yml,ecosia.yml,kagi.yml,yandex.yml, andyahoo-japan.yml. The GitHub HTML directory view could not be retrieved in this environment; the API endpoint worked.yahoo-japan.ymlis not a match for Anubis's supported Yahoo engine, and there is no Mojeek file, so those remain explicitly unwatched rather than guessed. - Shipped experiment: move the mapping into
.github/engine-watch.json; test that every engine inutils/engines.tsis either watched exactly once or has a stated reason it is not. The watcher also fails if a mapped upstream path disappears. Add local-repository tests for recent upstream changes, deduplication, unrelated files, old commits, and missing mapped files. Remove the issue-watch's silent failure and raise its pagination limit. - Ran against upstream: cloned
ublacklist/builtin, collected the existing engine-issue bodies withgh, and ran the watcher without creating an issue. It found two Google commits from 2026-09-27: desktop video/gotolinks now fall back to the displayed domain, and image links can fall back todata-lpage. Anubis already uses the displayedcitefallback for opaque links on supported web results; image-search pages are intentionally excluded, so neither change needed a selector change here. - Rejected: generating Anubis selectors directly from SERPINFO. The upstream rules describe a different matching model and don't cover Anubis's structural boundaries, page cleanup, or pagination; generated selectors could appear current while breaking real pages. The weekly issue remains a signal for human review and live-page confirmation.
- From an issue to a pull request (2026-09-29): an issue only linked to upstream commits, so each one meant reading uBlacklist's history by hand, and nothing in the repository recorded which version of its rules Anubis was last compared with. Anubis now keeps a copy of the eight watched files in
upstream/serpinfo/(MIT, licence alongside), and the weekly workflow updates the copy on a branch and opens a pull request. The review is the diff;source.jsonrecords the upstream commit, so the next sync lists only the commits since. Still not generated selectors (rejected above): the copy is data for review and tests, and the build doesn't read it. - Checking addresses against the copy: a test reads every address in each copied file's
matchesand checks the engine's ownmatchescovers it. Run against uBlacklist atb2f1acc(2026-09-27), it found Bing'swww2andwww4hosts and 14 Yandex country domains (yandex.az,.by,.co.il,.com.am,.com.ge,.ee,.eu,.kz,.lt,.lv,.md,.tj,.tm,.uz) that Anubis didn't run on; they were added. Google's country list and DuckDuckGo's hosts already matched. A sync pull request now fails CI when uBlacklist starts matching a new address, which is the one kind of upstream change a test can judge by itself. Adding hosts widens the content script's matches, which is fine before the first store release; after it, Chrome asks users to approve new sites, so weigh each addition then. - CI on a workflow's pull request: pull requests opened with
GITHUB_TOKENdon't start other workflows, so the required check would never run. The sync startsci.ymlon the branch withworkflow_dispatch, which is allowed, and the check reports on the branch's commit. Not verified on GitHub yet: the first scheduled run will show whether the protected branch accepts that check. - uBlacklist as a second opinion on the mocks: every search-page check runs against mocks in
e2e/fixtures.mjs, written without access to the real engines, so nothing said whether a mock still resembled the engine. uBlacklist's result selectors are kept current against the real pages, sotests/mock-realism.test.tsparses each mock (with linkedom, no browser) and counts the elements uBlacklist's web-result selector finds. On 2026-09-29 all six (Google desktop, its awkward layout, its phone layout, DuckDuckGo, Bing, and Brave) found exactly one per result, which is the first independent evidence that the mocks use the real class names. The test also checks each selector still appears in the copied rules, so when the weekly sync brings a new one, the pull request names the mock to update. It checks result containers only: uBlacklist's rules say nothing about Anubis's headings, boundaries, clean-up, or paging.
Engine watch
- A weekly workflow opened an issue when uBlacklist changed its rules for an engine Anubis supports (replaced by a pull request on 2026-09-29, below). The development sandbox can't reach GitHub's API, so the script reads upstream history with
git logover a blobless clone and only files the issue throughgh. Run against real history, it reported four Google changes in 45 days; commits already mentioned in anenginesissue are skipped, so the ten-day window can overlap safely.
Bugs found while testing
- Show on one hidden result reverted when the mouse moved. The button set
data-anubis-revealon the result directly. Engines rewrite parts of the page on hover, that runs another pass, and the pass reset the attribute from the page-wide "Show hidden" state. Results shown one at a time are now remembered by URL until the next search. The e2erevealpart clicks Show, changes the page and moves the mouse; it failed on the previous build. - The popup got no numbers from the page in Chrome. The content script replied to messages by returning a promise, which Firefox accepts and Chrome ignores, so the popup's "This page" section never filled in Chrome. Both scripts now reply with
sendResponse. Found because the e2e harness asked the page for its stats and gotundefined. - Testing closed shadow roots. Page scripts and Playwright locators can't reach into closed shadow roots, so e2e clicks buttons there through the DevTools protocol (
DOM.getDocumentwithpierce: true, then the button's box).
Google in Firefox: flipped tags, missing buttons, misplaced summary
Follow-up: after the fixes below, the summary still sometimes appeared below the first results. The anchor was the first result in the list holding most results, and Google nests some results one level deeper than the rest: a first result with sitelinks, or a group of results. The first result then wasn't in that list, so the summary went below it; with six sitelinks, the sitelinks became that list and the summary went inside the first result. A grouped Google mock (six sitelinks under the first result, two results in a group) reproduces this on the previous build.
- Summary anchor, again. The results area is now the list holding most results, widened to the engine's boundary (Google's
#rso). The summary goes above the first result anywhere in that area, as a direct child of the smallest element holding every result in it. - Sitelinks. Each sitelink has its own
h3, so each was weighed as a result, with its own tags and weigh button, and it cut the first result's container short (its parent held a second heading), leaving the snippet and sitelinks outside it. A heading that links to the same site, and shows no address of its own, now counts as part of the result above it. Grouped results each show an address, so they stay separate. - To confirm live: that Google's sitelinks still use
h3inside a link and show no<cite>, and that#rsostill holds the results.
Reported from real use: on Google in Firefox the "Reference" tag and "Hidden" read upside down, the weigh button appeared on only the first two results, and on page 2 the summary sat under the first result with no weigh buttons at all. Google couldn't be loaded from the development sandbox, so the fixes target every mechanism that produces these symptoms, and a hostile Google mock (google(…, hostile = true) in e2e/fixtures.mjs) reproduces all three on the previous build and passes on the new one.
- Flipped text. Some layouts flip a wrapper with a transform and flip their own children back, which leaves anything else inside the wrapper upside down. The tag row was inserted inside the title link's wrapper. It now climbs out of wrappers that hold only the title, steps outside any transformed wrapper, and, as a last resort, adds up its ancestors' transforms and applies the inverse so its text reads upright.
- Page CSS reaching Anubis's elements. A shadow root protects its contents, not the host element, and the page's stylesheets can still match the host (
… > :last-childand the like), hiding or transforming it. Every host now pins display, visibility, opacity, position, and all transform properties inline with!important, which beats any page rule. - Nesting depth. Finding a result's container gave up after 8 levels, a limit carried over from the original script. Google already nests a result about 8 deep, so a few more wrappers left Anubis on an inner block: the weigh button, hiding, and reranking all acted on part of a result, and the summary landed inside the first result. The limit is now 30; the real stops are the results boundary and a parent holding a second result.
- Summary anchor. The summary went before the first heading in page order, which can belong to something else. It now goes before the first result in the list that holds most of the results (revised in the follow-up above).
- The weigh button now rests at low opacity instead of being invisible until hover, as the original block icon was always visible.