{"title":"The instrument, published before the data","what_this_document_is":"The measurement, written down before there is anything to measure. It says what counts as use, what gets logged, how long it is kept, and what will be published — so that a report six months from now cannot be accused of choosing its measure after seeing the numbers. If you connect an agent to this endpoint, this is what happens to that call.","question":"Does anyone actually connect an agent to a personal career site, and what do they ask it?","document_status":{"state":"pre-launch","last_updated":"2026-08-08","meaning_of_pre_launch":"The endpoint is live but unannounced, and measurement_start is not set. Everything here is being settled BEFORE the question is asked of anyone.","commitment":"Once measurement_start is set, every subsequent change to this document is added to `revisions` below, with the date and the reason. A pre-declared measure that can be quietly edited afterwards is not a pre-declared measure. Changes to the DEFINITIONS of what counts will not be made after that point at all; only clarifications, corrections of fact, and additions.","revisions":["2026-08-06 — threshold, demand classes and reporting commitment set (Rafe).","2026-08-07 — traffic type and acquisition separated into two dimensions; they had been collapsed into one mutually exclusive class, which lost information (an invited ChatGPT user is both human-invoked and invited, and only the invitation was recorded).","2026-08-07 — raw query retention raised from 30 to 180 days; see kept_180_days.","2026-08-07 — the permanent term ledger narrowed to counted traffic only.","2026-08-07 — generic client libraries reclassified from human_invoked to external.","2026-08-07 — rewritten for a reader rather than a maintainer; three internal inconsistencies fixed, listed under corrections.","2026-08-07 — landing-page analytics added (separate Umami site ID, cookieless, page views excluded from the threshold); WebMCP tools noted as present on rafeblandford.com behind an experimental browser API.","2026-08-08 — corpus revision added to every response and every logged call (follow-up brief §4.2, brought forward from post-launch). The corpus is regenerated daily, so without it no trace could say which evidence answered it. Brought forward rather than deferred because a trace written without the revision can never gain one afterwards.","2026-08-08 — acquisition split from two kinds into three: invited, channel, unknown. A `via` token had meant only 'Rafe invited this person', so a token embedded in a public listing would have been reported as an invitation and subtracted from organic discovery — the exact opposite of what it evidences. Made before any tagged link was published and before measurement_start, which is the only point at which a definition here can still change.","2026-08-08 — connections.ndjson added: handshakes are now recorded, in their own file, and remain excluded from the threshold. This ADDS a signal rather than changing what counts, and it is being made before measurement_start is set. The reason is that a signed-in ChatGPT session on 8 Aug connected and enumerated all four tools while producing zero log lines, which would have left a six-month null unable to say whether anyone had arrived at all. See connection_counter."]},"pre_declared_threshold":{"number":20,"definition":"Genuine externally initiated MCP calls in six months — the aggregate of the counted traffic types (human_invoked + external), invited or not — excluding Rafe's own tests, uptime monitors and automated crawlers. The aggregate decides whether the threshold was crossed; every published figure also shows the split.","if_below":"The honest answer is 'essentially nobody used it', and that gets published.","if_above":"The counts get published anyway, split by class.","set_by":"Rafe, 6 Aug 2026, before launch."},"demand_hierarchy":{"note":"Two dimensions, not one list. What kind of caller it was, and how they found it. Kept separate because an invited ChatGPT user is both, and collapsing them loses the caller type.","traffic_type":{"note":"What kind of caller. Recorded as `demand_class`.","self_test":"Rafe's own probes — probe token, tailnet, or a rafe-smoke agent. Excluded.","monitor":"Uptime checks. Excluded.","crawler":"Automated collection — *Bot and *-SearchBot agents building a training corpus or a retrieval index. Excluded.","human_invoked":"COUNTS. A person driving an agent, identified by an agent string in which the vendor asserts a human — Claude-User, ChatGPT-User and the like. Anthropic's own MCP connector presents as Claude-User.","external":"COUNTS. Everything else, including generic client libraries such as openai-mcp and modelcontextprotocol. Those name software, not a person: the same library serves a signed-in chat session, a developer tool, an API integration and an automated evaluation. Attribution stays conservative on purpose, and it costs the measurement nothing, because external counts too."},"acquisition":{"note":"How the caller found it. Recorded as `via` — a token on the connection link. Orthogonal to traffic type: an invited ChatGPT user is both human-invoked and invited. THREE kinds, because a token means one of two opposite things and its absence means a third.","invited":"A token Rafe handed to a named person. Closer to a functional test with feedback attached than to demand, and every published figure separates it out, because the total will be primed by invitations.","channel":"A token embedded in something Rafe published rather than sent to anyone — the MCP Registry entry, the launch article, the LinkedIn post. THE PERSON FOUND IT THEMSELVES; the token only records which route they took. This is organic discovery and it is the finding the experiment is looking for, so it is never reported as an invitation.","unknown":"No token. Also organic, with nothing to say about the route — and this is expected to be the LARGEST group rather than a residue, because the canonical address is published untagged on the landing page and that is the route the endpoint is designed around: someone reads the site, decides to connect, and copies the address. Acquisition can only distinguish routes that publish their own distinct link. It should not be read as 'found it with no help'.","declared_channel_tokens":["registry","plugins-directory","article","linkedin","card"],"why_the_channel_list_is_published":"So that 'which of these is a channel rather than an invitation' is auditable rather than a naming convention. Invitation tokens are deliberately NOT published: they are people's names.","acquisition_is_asserted_not_verified":"Acquisition is asserted, not verified — like client identity. Channel tokens are public by construction — printed in a registry entry and in an article — so anyone can append one, and a forged channel token would inflate the ORGANIC figure. That is the flattering direction, where a forged invitation token could only ever deflate it. It is reported as a claim rather than treated as authenticated, because there is no version of a public listing whose token is also a secret.","corrected_2026_08_08":"Until this date the report counted ANY token as an invitation, and printed 'of which invited N, the rest found it themselves'. A registry pickup would have been subtracted from the number it is strongest evidence for. Found before any tagged link was published and before the clock started."},"how_the_threshold_uses_them":"The threshold is crossed on the aggregate of the counted traffic types. That aggregate exists only to decide whether the line was crossed. Every published figure also shows the split — by traffic type, and by all three acquisition kinds — because the total will be primed by direct invitations, and 'someone asked Claude about Rafe', 'something connected to this on its own' and 'Rafe asked a person to try it' are different findings.","what_is_not_demand":"A request that never reached a tool. A crawler issuing GET on the endpoint gets a 405, a bad Origin gets a 403, and an over-rate caller gets a 429. All are logged, because failures should be visible, and none is counted as use. Only tools/call and resources/read count.","connecting_is_not_demand_either":"A client that connects and asks what is here — tools/list and the other enumeration methods — has not used anything. Those handshakes are recorded, in their own file, and are never counted towards the threshold. See connection_counter under what_gets_logged for why they are recorded at all.","identity_is_asserted_not_verified":"Client identity is asserted, not verified. There is no authentication here, so anything can claim to be any client, and every client name in this measurement is a claim rather than a fact. It is recorded as a claim and will be reported as one."},"what_gets_logged":{"per_request":"Timestamp, the method and tool or resource used, latency, how many results came back, the error type if it failed, which page URLs were returned as evidence, the traffic type, the acquisition token if the link carried one, the client name and protocol version the caller asserts, and which version of the corpus answered (see corpus_revision).","corpus_revision":"Every response carries a `corpus` block — a contract version, the dossier version, the generation timestamp and a SHA-256 digest of the public corpus — and the digest is recorded against every logged call. The reason is that the published corpus is regenerated daily, so six months is on the order of 180 distinct versions of it: without this, a report could say that a question was asked and nothing about what the answer was drawn from. It also means anyone who records an answer from this endpoint can prove later which corpus produced it. The digest deliberately EXCLUDES the generation timestamp, so a regeneration that changed nothing produces the same digest; identical content always gives an identical digest, and any changed published claim changes it. Added 8 Aug 2026, before the measurement window opened, because it cannot be applied to traces after the fact.","no_ip_address":"The caller's IP address is NOT stored. It is used in memory to apply rate limiting and to recognise Rafe's own network, and is not written to any file.","about_query_text":"For search_writing and check_experience the long-lived file records a hash of the query and the NUMBER of terms in it — not the words. The words go to a separate, shorter-lived file described below.","kept_indefinitely":"Three files. calls.ndjson, one line per request as described above, with no free text at all — this is the six-month instrument. topics.ndjson, the normalised terms of queries from counted traffic only (see permanent_topic_ledger), which is the record of what people looked for and did not find. And connections.ndjson, one line per handshake (see connection_counter).","connection_counter":"When a client connects and enumerates what is here — tools/list and the other listing methods — that is recorded in connections.ndjson: timestamp, which method, latency, the traffic type, and the client name and protocol version the caller asserts. Same fields as a call, no query text, no IP address. It is NOT demand and is never counted towards the pre-declared 20; it is kept in a separate file so that it cannot be added to the count by accident. WHY IT IS RECORDED AT ALL: without it, a client could connect, read the full tool list and call nothing, leaving no trace whatsoever — so a six-month result of zero could not distinguish 'nobody connected' from 'people connected, looked, and found nothing worth calling'. Those are different answers to the question this endpoint exists to ask, and the second is the more interesting one. WHAT IT IS NOT: it counts handshakes, not people. Sessions here are stateless, so one client reconnecting five times writes five lines, and a client that re-establishes per conversation writes one per conversation. It is an upper bound on curiosity, and any published figure will say so. Added 8 Aug 2026, before the measurement window opened.","kept_180_days":"queries-YYYY-MM-DD.ndjson — the query words themselves, with contact details stripped: email addresses, phone numbers, URLs, long digit strings and @handles. Raised from 30 days on 7 Aug 2026, because the six-month report is written at around 180 days and a 30-day file would have expired its own evidence. Still bounded rather than indefinite, and never published without consent.","what_redaction_is_not":"Pattern matching, not entity recognition, and not anonymisation. It will not remove a person's name, a company name, or a job description that identifies someone by role plus employer. Please do not paste anything into a query that you would not want kept.","permanent_topic_ledger":"The normalised terms of each query are kept indefinitely, and ONLY for counted traffic types (human_invoked and external). The argument for keeping them forever is that a stranger's unanswered question is a content backlog worth having; that argument does not cover Rafe's own probes, an uptime monitor or a crawler, so those are never written to it. For a one-word query the normalised term is the word that was typed, so treat it as retained rather than discarded.","never":"Anything in the private source dossier that is not marked public. The server process does not hold that file at all; it reads only generated public projections of it.","rate_limit":"60 requests per 60 seconds per caller and method, as a backstop. Exceeding it returns 429 and is logged as a rejection, not as use.","rotation_at_announcement":"calls.ndjson, connections.ndjson, topics.ndjson and every pre-launch query file are rotated together, in one operation, at the moment measurement_start is set. Otherwise the permanent backlog would open with Rafe's own test vocabulary while the traffic number opened clean.","landing_page_analytics":"The human landing page at / carries Umami — self-hosted, cookieless — on its own site ID, separate from rafeblandford.com's. It is NOT on /mcp, and page views are NOT part of this measurement: they never count towards the pre-declared threshold, which counts MCP calls only. The reason for a separate ID is that the people who read the page and the agents that call the endpoint are different populations, and merging them would lose both. Added 7 Aug 2026 and disclosed on the page in the same change."},"reporting":{"checks":["48 hours — technical failures","one month — early behaviour","monthly thereafter"],"report":"A field report at six months, not a benchmark and not a study: descriptive counts, privacy-safe item-level outcomes, a failure taxonomy with concrete examples, changes made because of not_found, and an explicit null if no demand appeared.","commitment":"Published either way.","uptime":"An uptime monitor checks the endpoint independently, so that a quiet outage cannot be mistaken for an absence of demand. Its calls classify as monitor and are excluded."},"constraints":{"maintenance_ceiling":"Under one hour a month after the first month. If it exceeds that, it fails its own test.","reversibility":"One Caddy block and one container. Removing it removes it.","human_site":"Unaffected. Nothing in this experiment changes how rafeblandford.com behaves for a person."},"what_this_is_not":["Evidence that machine demand exists. That is the hypothesis, not the premise.","A controlled study. Model versions, routing and client behaviour all move underneath it.","A complete record of Rafe's work. It serves what the site publishes, which is a selection.","A verified identity check on callers. See identity_is_asserted_not_verified.","A statement about Rafe's experience. check_experience answers only from what this site publishes, which is a selection of his work; not_found means the corpus does not evidence a topic, never that he lacks it."],"corrections":{"note":"Things this instrument got wrong before launch, kept rather than quietly fixed. Each would have biased the result, and all were found and corrected before the measurement window opened.","2026-08-07":["Claude-User was classified as a crawler and excluded, because the crawler list was copied wholesale from the website's own analytics. That is right for a web page and wrong for this endpoint: on an MCP endpoint a Claude-User call IS a person driving an agent. Left uncorrected, the six-month number would have been structurally zero regardless of genuine use.","Rafe's own model-driven evaluations reach the endpoint through the same connector, and so were indistinguishable from real demand. They now tag themselves with a probe token in the URL and classify as self_test.","Transport-level rejections were counted as demand. Six phantom calls had accrued against the threshold before it was caught.","Six calls from a generic OpenAI client library were briefly classified as human-invoked. The client string names software, not a person, and the calls did not support the stronger reading. Reverted to external, which still counts.","This document said the permanent term ledger held the terms of EVERY query while another section correctly said counted traffic only. The narrower statement was the true one; the contradiction is fixed here."]},"what_is_served":{"tools":["get_evidence","get_post","search_writing","check_experience"],"resources":["rafe://profile","rafe://career","rafe://case-studies","rafe://provenance"],"resource_templates":["rafe://writing/{slug}"],"prompts":["assess_role_fit","brief_me","what_does_he_think_about","interrogate_the_evidence"],"mutating_tools":0},"how_to_challenge_this":"If something here is wrong, or you want the record of your own queries removed, the contact route is https://machines.rafeblandford.com/.well-known/security.txt — it is a person, not a form, and there is no automated route to Rafe from this endpoint by design.","measurement_start":{"value":null,"note":"Set this to the exact ISO timestamp of the ANNOUNCEMENT, not the deployment. The endpoint went live 7 Aug 2026 and sat unadvertised; 'ships' means the launch article and the invitations, because that is when the question 'does anyone use this' starts being asked. Every series is rotated in the same operation and under the same timestamp — see rotation_at_announcement. This said 'rotate calls.ndjson' until 8 Aug 2026, which contradicted that section and would have left the permanent term ledger opening with pre-launch test vocabulary while the traffic number opened clean."}}