Publish a privacy policy and link it from every page - #140
Merged
Conversation
The Anthropic Software Directory Policy requires a clear, accessible privacy policy link, and the site had no such page. The site's footers already stated the position accurately, so this elevates an existing statement into a page rather than inventing a position. The page states what the site does and does not do, verified against the code: no cookies from the site's own code, no client-side analytics, no browser storage, no forms, and no third-party embedded resources. It then says plainly what Cloudflare and GitHub see, because a flat claim of zero tracking would be false. It also covers the privacy mailbox itself, which is the one channel where a reader chooses to hand over personal data. Cloudflare's privacy policy and GitHub's privacy statement are linked; both URLs were fetched and confirmed before use. The footer navigation gains a Privacy link on all eleven pages that carry it, including 404.html. The ten dated pages have their article:modified_time and JSON-LD dateModified moved to today, because editing them otherwise leaves a stamp older than the commit and fails the metadata gate. The sitemap is regenerated and now lists eleven pages. Open for confirmation before merge: the page states that no data protection officer is appointed, which is a claim about the maintainer's arrangements rather than something readable from the repository.
Deploying cleanlanguageai with
|
| Latest commit: |
615444a
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://b4bdb67b.ai-language.pages.dev |
| Branch Preview URL: | https://privacy-policy.ai-language.pages.dev |
Codex and gemini both returned DO-NOT-SHIP. Every finding below was validated at source before it was applied. Accuracy fixes. The page claimed GitHub serves the downloads; the install page serves 3 Markdown instruction files from this site and only the archive from GitHub. It listed posluns.ca as an outbound link destination, which it is not, since that domain appears only in JSON-LD, and it omitted the Cloudflare and GitHub Docs links added in the same commit. It called the skill a text file, when the package is 5 Markdown files, an icon pair, and a configuration file. It said the traffic counts describe pages rather than people, when Cloudflare derives a unique-visitor figure from IP addresses. It said the 2 scripts store nothing, when the copy button writes to the clipboard on click. Completeness fixes. The privacy mailbox now discloses Cloudflare Email Routing and the destination mailbox provider. GitHub issues are disclosed as a public channel. The rights section now routes requests to the maintainer, because Cloudflare directs visitors back to the site operator rather than answering them. Overclaiming fixes. The page no longer asserts a lawful basis, no longer implies that published provider policies stand in for transfer agreements, and no longer pairs "we do not share" with operational disclosures that contradict it. House style. The site is a public example of its own standard, and the draft broke it: throat-clearing openers and meta-commentary are removed, and quantities now use numerals, per SKILL.md line 50 and the removal list at lines 72 and 77. Also fixes the new file's permissions, which were 0640 against 0664 on every sibling page. The committed git mode was always 100644, so the live site was never affected, but cross-family reviewers could not read the file.
Codex and gemini both returned DO-NOT-SHIP again. Ten findings, each validated at source. Claude could not review either round: the pool reported no logged-in worker, so this converged on 2 families rather than 3. The page had claimed more certainty than it could support. It said visitors send us nothing, then described the request data their browsers send. It said the data that could be demanded sits with other companies, while also saying we hold privacy correspondence. It said a user's text stays between them and their AI, which is a claim about providers we do not control. It said GitHub answers for issues, when issue content also reaches the maintainer as repository owner. It said we obtain no data from anyone else, when Cloudflare supplies derived analytics and GitHub sends issue notifications. Narrower corrections: the outbound-link list is now scoped to clickable links, since JSON-LD and a copied prompt also carry external URLs; the complaint route no longer assumes every visitor has an authority with jurisdiction; the mailbox is described as held with a third-party provider; and the "What this page does not cover" heading, which contradicted its own contents and read as meta-commentary under references/anti-patterns.md, is now "Other privacy points". The rights section no longer promises to approach Cloudflare on a visitor's behalf. Gemini was right that this is not a result the maintainer can deliver, and codex reached the same conclusion independently. Two items still need the maintainer: whether a data protection officer is appointed, and whether to name the mailbox provider.
The maintainer confirmed both. Jeff Posluns is the appointed data protection officer, so the page no longer says none is appointed. The privacy address does not forward to a separate email provider: Cloudflare Email Routing delivers it to a Cloudflare Worker behind the privacy and data protection interface he runs across his projects. The recipient list and the retention section follow that correction.
The maintainer confirmed that messages stay in the privacy interface until he deletes them, so the page states that rather than implying a fixed period. Round 3 reached only gemini: the claude pool reported no logged-in worker again, and codex returned an empty body on every account it tried. Of its three findings, two are applied. The outbound-link list is now non-exhaustive, because an exhaustive one goes stale the moment a link is added anywhere on the site. The GitHub paragraph now covers the content delivery servers that release downloads redirect to; gemini claimed the hrefs point there directly, which is wrong, since they are github.com URLs, but the redirect is real. Its third finding, that naming an appointed data protection officer asserts a formal legal designation the project may not want, is held for the maintainer. He stated the appointment as fact, so it is not the assistant's to reverse.
The maintainer accepted gemini's round 3 finding. Naming an appointed data protection officer asserts a statutory designation under GDPR Articles 37 to 39, with duties and liabilities attached, which a free project has no reason to claim. The page now says he handles data protection and privacy enquiries himself rather than through a separately designated officer, which satisfies the template's controller block without asserting the legal status.
I reported codex round 3 as a family failure. It was not: the worker completed with a PASS and a full report, but orch-verify wrote nothing to stdout and only persisted the deliverable to disk. Reported to lab_infra. Recovered and applied here. Accuracy. The page said Cloudflare analytics were the only measurement we have; GitHub also reports repository traffic and a download count per release asset. It said email is the one place people give us personal data, while also saying issue content reaches the maintainer. It covered GitHub issues but not pull requests, comments, reviews, or commit authorship, which reach the repository owner the same way. It said the site asks nobody for personal data, while asking readers to write to the privacy address. Scope. Responsibility is narrowed to how Clean Language handles personal data, rather than any personal data the site involves, since GitHub is its own controller for its platform. The notification claim is replaced with what is invariant: the maintainer can read public contributions as repository owner. Also removes a repetition I introduced in round 2, and adds the privacy page to the site/README.md contents inventory, which codex noticed was stale.
Ten findings, each validated at source. Gemini returned SHIP on the same commit; claude was quota-exhausted until 16:40Z and did not review. The page metadata said the site sets no cookies, dropping the qualification the body and footer both keep, so it contradicted them. It said we run no analytics of our own while also saying Cloudflare's are the only measurement we have. It said the scripts write nothing to browser storage, but install.js calls history.replaceState twice. It said the site collects nothing directly, which the privacy mailbox contradicts. The largest correction: we hold local clones of the public repository, and those carry contributors' names, email addresses, and commit dates. Codex found a real instance, commit 394a61f on the local pr133-check branch. The page now discloses that, and that git history persists in clones after something is removed from GitHub. Also: Canadian residence does not settle which regulator is competent, since Alberta, British Columbia, and Quebec have their own private-sector laws; the skill inventory was wrong, since the archive carries 2 icons plus licence and notice files; LinkedIn is another route to the maintainer, so the exhaustive channel claim is replaced with what the repository proves.
Claude round 5 returned SHIP with one LOW finding: 'no consent banner is needed' asserts a legal conclusion in the register of a fact, unlike every sentence around it. Substantively defensible, since the only cookie is a security cookie and the analytics are server-side, but it is not the assistant's or the maintainer's determination to state. Now reports what the site does instead.
Round 5 verdicts: claude SHIP, gemini SHIP, codex DO-NOT-SHIP with 9 findings. Six applied here, one was already fixed, and two were unverified items rather than defects. The page said correspondence was the only thing we hold, which its own local clones disclosure had already contradicted. It said you never hand the site anything to store, and that nobody is required to provide personal data to read or download, without preserving the qualification that reading still discloses a request. It said public GitHub contributions stay in the repository, when issues and comments can be edited or deleted; what cannot be recalled is a commit already pulled into someone else's clone. The copy button falls back to selecting text when the clipboard write fails. New disclosure: the Cloudflare Pages Git integration reads the repository to build the site, so Cloudflare receives repository contents and deployment metadata. That processing was undescribed. Codex's largest unverified item is now resolved empirically rather than left open. Fetching production returns no Cloudflare Web Analytics beacon, no external script tags at all, and no Set-Cookie on a plain request, which is what the page claims.
Round 6: claude SHIP, gemini SHIP, codex DO-NOT-SHIP with 9 findings. Eight applied. The material one is a use path the policy never covered. site/demo/index.html tells visitors to paste their writing into an AI without installing anything, and the prompt it gives them asks that AI to fetch the standard from GitHub. The page described only the installed skill, so it now covers the paste path and the GitHub request it can trigger. Also newly disclosed: GitHub Actions runs the checks on this repository and processes pushed content, commit and actor details, and run logs. Smaller corrections: public GitHub contributions are public by design rather than a legally compelled disclosure; Cloudflare's bot detection scores likely automated traffic rather than separating human from automated reliably; the 'what we receive from others' list excluded correspondence and contributions; two unsupported 'as on any website' generalizations are gone, one of which also read as sending Cloudflare data to other sites; the script count is replaced, since nothing gates it against the actual inventory; and the postal address rationale asserted things not in evidence. Not applied: codex wants defined retention purposes beyond answering the enquiry. The maintainer's stated criterion is that messages stay until he deletes them, and inventing further purposes would be fabrication.
All three families returned DO-NOT-SHIP on 89084cc, a regression from round 6. The cause was my own round 6 edits, not new surface. Claude and gemini independently caught the same grammar fault: replacing the script count with 'The site's JavaScript' left the verbs plural against a singular subject, on the page of a site that publishes a writing standard. Codex caught the knock-on effect of the two disclosures round 6 added. Having disclosed GitHub Actions run logs and Cloudflare Pages deployment records, the page still said we neither control, search, nor delete any of it. That is true of provider request logs and false of records inside our own accounts. Both the retention and rights sections now separate the two.
Ran claude-fable-5, gpt-6-astra at xhigh, and gemini-3.1-pro. Gemini returned SHIP; the other two returned DO-NOT-SHIP with no HIGH findings, the first round in this cycle without one. Accuracy. The page said a prompt asking an AI to fetch the standard sends that request to GitHub. The Copilot setup at site/install/index.html:228 points the agent at a copy on this site instead, so the request goes wherever the prompt points it. It also said Cloudflare's filtering scores each request, when Cloudflare documents requests that receive no score, and it warranted on Cloudflare's behalf that the filtering makes no significant decision; that is now stated as our own view. The claim that we can neither search nor delete provider logs was too absolute, since the Cloudflare dashboard exposes sampled security events. Density. The contact address appeared 3 times in 4 sentences, the Cloudflare bullet described the mailbox twice, and a sentence about the skill sending us nothing repeated one 2 sentences earlier. A hedge also understated what every request carries: an IP address, the requested address, and a time are always present, while the query and headers are conditional.
Six findings applied. Gemini's HIGH was that the sentence added in round 8 to attribute Cloudflare's bot filtering still stated a legal classification, just in our own voice. The page now describes what the filtering does and what its outcome is, and states our own practice, without classifying Cloudflare's processing. Codex: the list of what we receive from providers omitted the sampled security events the page elsewhere says we can inspect; the paste-a-prompt path assumed a hosted assistant, when the install guide supports local models, where nothing leaves the reader's machine; a request does not carry its own generation time and a cached page reaches no server at all; a 201-word paragraph in the rights section is split; and the Cloudflare bullet described the email routing twice. On that last one I first reported it refuted, having checked with a truncated grep that cut off the second mention. It was real. Corrected here.
Gemini round 10, HIGH: the page said following a download or source link makes a request to GitHub, but the Markdown instruction downloads on the install and instructions pages are served from this site. Only the archive and source links go to GitHub. Codex returned DEGENERATE this round and did not review.
Gemini SHIP. Claude and codex each raised one HIGH, and both trace to edits I made in round 9. Claude: adding the sampled security events made 'the only measurement we have' contradict two later sentences. Now qualified. Codex: 'a local model means nobody outside your own machine' is a guarantee I had no basis for, since a local assistant can still have tools and network access. It now says disclosure depends on the assistant, its configuration, and what it connects to. Codex also showed that blocking is an outcome of Cloudflare's filtering, not just challenge or serve. Also attributes two claims about the Cloudflare cookie to Cloudflare rather than stating them in our own voice.
Gemini SHIP. No HIGH findings from any family, the first round of the cycle with none. Applied: the sentence covering both analytics and security events lumped them together as counts, when the events show individual addresses; the GitHub sentence had a plural relative clause followed by a singular possessive reaching back two clauses; the international-processing section covered request data but not mailbox messages, which also sit on Cloudflare's global infrastructure; and the hero read as absolute against the mailbox section. Two of claude's findings quoted text replaced in round 11 and were refuted against the live file: its 'only measurement' contradiction and its double-colon sentence. Verified before discarding, not assumed. Left for the maintainer: whether 'we would disclose data only where the law required it' is a commitment he can keep, since it forecloses disclosures that are permitted rather than required, such as consulting a professional adviser.
Claude and codex independently raised the same HIGH and MEDIUM against line 120, and my round 12 edit caused it. Extending the international-processing sentence to cover mailbox messages made the next sentence, 'We hold no copy of that data ourselves', false: the messages are precisely what we do hold. A reader would have concluded an access request about their own message was pointless, which the rights section promises the opposite of. The sentence now scopes the no-copy claim to the request logs and names the messages as the exception. Round 13 was collected by binding each result to the deliverable path its dispatch reported, rather than by timestamp window. Claude's report cites 01f9124, confirming it reviewed the intended revision.
The maintainer gave the accurate position: messages are not kept beyond reading and answering them, and other platforms do whatever is default for them. The page previously said messages stay in the privacy interface until he deletes them, with no fixed retention period, which came from an earlier and less precise statement. Five places carried the old position and now carry the new one: the mailbox section, the retention section, the disclosure paragraph, the international-processing sentence, and the rights section, which had promised to act on correspondence we no longer keep. Not changed: the local clones disclosure. 'We have no data' holds for correspondence and provider logs, but local copies of the public repository do carry contributors' names, email addresses, and commit dates in git history, verified against a live commit. That stays as fact.
The maintainer's correction: the names and addresses in git history are the ones contributors themselves publish through GitHub's pull-request process, public by GitHub's design, and nothing to do with a collection decision of ours. The page had framed local clones as a store of contributor personal data we hold, which put our own choices where GitHub's process actually sits. It now says what is true and relevant: public contributions stay where GitHub keeps them, commit authorship is what the contributor published, and git history is distributed by design, so it travels with every copy anyone makes, ours included. The page makes no claim either way about whether that is personal data, which is not its question to settle.
Claude SHIP, gemini SHIP, codex DO-NOT-SHIP with no HIGH findings. Codex: Cloudflare's email routing records outlive the message they describe and we can inspect them, which the retention section omitted while the rights section said we keep little to act on. Now stated. Removed the disclosure sentence codex flagged, 'Beyond that, we would disclose data only where the law required it'. My own commit message had left it open for the maintainer, and codex was right that an unresolved commitment should not stay published. The paragraph still rules out selling, advertising disclosure, and data brokers, and says there is little to disclose. Claude: a cache anywhere between reader and site can serve a page, not only the browser's own; 'sets' fits cookies but not Web Storage or IndexedDB; and the two public-contribution sentences carried the same fact, with a bare 'address' that could read as the postal address discussed earlier.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The Anthropic Software Directory Policy requires a clear, accessible privacy policy link, and the site had no such page. This is tracked as item 27.17 in the durable store.
The site's footers already stated the position accurately, so this elevates an existing statement into a page rather than inventing a position. Every claim on the page was verified against the code: no cookies from the site's own code, no client-side analytics, no browser storage, no forms, and no third-party embedded resources. The page then says plainly what Cloudflare and GitHub see, because a flat claim of zero tracking would be false.
It also covers the privacy mailbox itself, which is the one channel where a reader chooses to hand over personal data.
Cloudflare's privacy policy and GitHub's privacy statement are linked. Both URLs were fetched and confirmed before use.
The footer navigation gains a Privacy link on all eleven pages that carry it, including
404.html. The ten dated pages have theirarticle:modified_timeand JSON-LDdateModifiedmoved to today, because editing them otherwise leaves a stamp older than the commit and fails the metadata gate. The sitemap is regenerated and lists eleven pages.All ten local gates pass, each run to a real exit code.
Open for the maintainer before merge: the page states that no data protection officer is appointed. That is a claim about the maintainer's arrangements rather than something readable from the repository, so it needs confirmation.