- A crawler sensor — Cloudflare pull or edge middleware, one per hostname. This is the only sensor that sees AI crawls, answer fetches, and 404 demand.
- Website analytics and Search Console — connect GA4 or PostHog for referral sessions and conversions. Connect Search Console for Google search clicks and impressions.
- The tracking snippet — an optional first-party record of AI referral page views.
Where Setup Lives
Sensors are connected and repaired in Settings › Integrations. You can also reach it from Home: select Coverage, then Integrations. The tab groups connections into Crawler sensors, Website analytics, and Search performance. Each row shows its status and latest activity. Snippet and middleware rows show up to 20 recently observed hostnames from the current configuration. This is a recent sample, not a complete list of hosts or a guarantee that every route is instrumented. Choose Set up to connect a provider, or Manage to review an existing connection. Use one analytics provider per site: GA4 or PostHog.Connect a Crawler Sensor
AI crawlers and answer fetchers often do not run JavaScript, so DevTune reads them server-side. Choose one of two crawler sensors per hostname.Crawler Sensor Capabilities
Recommended split:
- Use Cloudflare pull for docs, marketing, or static hostnames that already sit behind Cloudflare.
- Use edge middleware for Vercel, Next.js, and app hostnames where you do not want a Cloudflare-to-app double proxy.
- Keep a hostname on one crawler sensor at a time. If you move a hostname from Cloudflare pull to middleware, remove or pause the Cloudflare hostname connection.
Cloudflare Pull
DevTune pulls hourly AI crawler analytics from Cloudflare’s GraphQL analytics API. Nothing is installed on your site, and your request volume never reaches DevTune — only aggregated hourly buckets do. The API token is stored encrypted and used only for readonly analytics pulls. When Cloudflare includes the sending address in those aggregates, DevTune checks it against the claimed operator and discards it before storing the bucket. If that address dimension is unavailable, the traffic remains visible and is marked as having no address from the sensor. Before you start, confirm the hostname you want to measure is proxied by Cloudflare — the orange cloud in your DNS records. Cloudflare has no analytics for a hostname it does not proxy, so a grey-cloud record returns nothing.Create a Read-Only API Token
In the Cloudflare dashboard, go to My Profile → API Tokens → Create Token, then Create Custom Token:
Nothing else is needed. Do not grant edit permissions — DevTune only reads.
You also need the Zone ID for that zone. It is on the zone’s Overview page in Cloudflare, in the API panel on the right.
Add the Source
In Settings › Integrations, open Cloudflare under Crawler sensors:- Under 1. Cloudflare source, paste the Zone ID and the API token.
- Give the source a Label so you can recognize it later, such as
example.com zone. If you leave it blank, DevTune names it after the zone ID. - Select Save source.
Add a Hostname
Once a source is saved, 2. Add hostname appears:- Select the Tracked domain this connection reports into. Only domains already tracked on the project appear here — add the domain in Settings first if it is missing.
- Check the Cloudflare hostname. DevTune prefills the tracked domain’s hostname; change it if the traffic you want arrives on a different hostname in the same zone, such as
docs.example.com. - Select Add hostname.
example.com and docs.example.com needs two.
What Happens Next
The sync runs hourly, on the hour, and reads a settled window of Cloudflare analytics — the most recent minutes are skipped because Cloudflare has not finished aggregating them. The first sync reaches back up to 24 hours, so a connection made this afternoon can arrive with yesterday evening’s crawls already in it. Subsequent syncs overlap the previous window so nothing falls through a gap. DevTune records each hourly bucket by page, user agent, bot class, and edge status, then classifies it against the AI bot registry. Because Cloudflare reports the real edge status for everything it proxies, missing pages reach 404 Demand with no extra work on your side. Edge middleware feeds the same view, but only for the routes where it can observe the 404 itself.Managing and Removing a Connection
Each connection row carries its own controls:- Disable stops the hourly pull and keeps the connection and its history. Enable resumes it.
- Delete removes the connection, and removes the stored credential when no other hostname is using that source.
- To change the hostname on an existing connection, select its tracked domain under Add hostname, edit the field, and select Update Cloudflare.
- To rotate the API token, select Add source, re-enter the same Zone ID with the new token, and save. DevTune replaces the stored credential and reactivates every hostname on that source.
If Cloudflare Stops Syncing
When Cloudflare rejects the token, the source reports Cloudflare token needs to be updated and syncing pauses for every hostname using it. Select Enter a new API token on that message and save a replacement. Syncing resumes on the next hourly run. For any other failure, the sensor reports Cloudflare sync needs attention. Check, in order:- The token still carries Zone → Analytics → Read for that zone.
- The zone ID matches the zone that serves the hostname.
- The hostname is still orange-cloud proxied.
- The connection is enabled rather than disabled.
Edge Middleware
Use the middleware package for application servers and edge runtimes. It forwards only matched AI bot requests, so normal browser traffic is never sent to DevTune.Generate the Ingest Key
In Settings › Integrations, open Edge middleware under Crawler sensors, select Generate ingest key, and copy the Server-side ingest key into a server or edge environment variable asDEVTUNE_AI_TRAFFIC_INGEST_KEY. Never expose it in browser code.
Install the Package
DEVTUNE_AI_TRAFFIC_ORIGIN to the application’s canonical public origin, such as https://www.example.com. Pinning the origin prevents client-supplied forwarded headers from changing the hostname recorded in DevTune. If one process serves multiple tracked hostnames, omit the fixed origin only when a trusted proxy overwrites X-Forwarded-Host and X-Forwarded-Proto, or create a separate adapter instance with its own origin for each hostname.
Next.js
For Next.js 16, addproxy.ts. For older Next.js projects, use the same body in middleware.ts and export middleware() instead.
Express
next() immediately and records the actual response status after Express emits finish.
Node HTTP
response.statusCode without blocking or changing the response.
Verifying Crawler Identity
A user agent is a claim, not an identity: anything on the internet can sendGPTBot/1.3. From @devtune/ai-traffic 0.3.0 the middleware sends the address
each matched request arrived from, so DevTune can check it against the
operator’s published networks and then discard it.
This needs
@devtune/ai-traffic 0.3.0 or later. Earlier versions send no
address, so their crawl and fetch records are stored as unverified. Each
checked request is stored with one of three verdicts: verified, failed, or
unknown when the check could not be made. Unknown requests say whether the
sensor supplied no address, DevTune has no identity check configured for the
operator, or an available check could not complete. Older unknown requests
appear as before verification. The hourly and daily summaries carry those
counts beside verified and failed totals, and the AI Traffic API exposes them
on recent answer fetches.Supported identity checks
DevTune checks request IP addresses against the operator’s published sources:
Mistral’s published feeds cover
MistralAI-User and MistralAI-Index, not its
training crawler. ByteDance has no configured identity check and remains unknown.
A platform name inferred from a user-agent or referral is not proof of bot identity.
Claude Code and Gemini CLI can fetch from the user’s own network. DevTune
recognises their user agents, but keeps their identity unknown unless an
independent operator check succeeds. An address outside hosted crawler ranges
is not an identity mismatch for these clients. OpenCode’s self-identifying
retry requests appear under Other and remain unknown. These clients’ fetches
count as answer fetches, not training crawls.
The Claude Code verification change runs in DevTune and needs no middleware
update. Middleware versions before 0.5.0 omit Amazon and DuckDuckGo registry
entries; update the package for that coverage. Cloudflare pulls need no package
update. From 0.5.0, middleware accepts new platforms and bot rules from DevTune’s
remote registry without further package upgrades. It refreshes in the background
on the first request and after the one-hour cache expires, including on browser
requests. Browser request data is not sent. Newly added bots may be missed until
the refresh completes; failed refreshes keep the last working rules. Address
extraction, adapter and event-format changes may still require an upgrade.
No proxy configuration change is needed if caller IPs already arrive.
Historical unknowns and mismatches remain unchanged because their IP addresses
were not retained.
The package only reads that address from a source your deployment declares as
trustworthy, because a forgeable one proves nothing. On Vercel it reads the
platform’s own forwarded header. On Express and Node it uses the socket peer.
X-Forwarded-For is only read when you select CloudFront or configure a numeric
offset, since a visitor can supply entries before the addresses appended by your proxies.
An edge runtime with no socket and no declared proxy sends no address, and that
traffic cannot be verified.
CloudFront
For CloudFront connecting directly to the app running the middleware, update to@devtune/ai-traffic 0.4.0 or later and add this option to your existing setup:
X-Forwarded-For; this setting reads the rightmost entry.
See AWS’s custom-origin documentation.
Use this setting only when requests reach the app through your trusted CloudFront
distribution. If an ALB or another proxy between CloudFront and the app appends
addresses, use a numeric trustedProxy offset instead: 0 selects the rightmost
entry, 1 skips one appended address, and so on. Confirm whether each proxy
appends, preserves, or replaces the header before choosing an offset.
This setup applies to middleware at the origin, not code running inside
CloudFront Functions or Lambda@Edge. After redeploying, future requests with a
usable address can be checked. Previously collected requests without an address
remain unknown.
Delivery and Statuses
The default client configuration starts a flush when 10 matched events queue or after 250 ms, and each ingest request drains up to 100 queued events. Express and Node servers benefit directly from batching. Next.js hands delayed sends toevent.waitUntil(), so responses are not delayed. If a short-lived runtime cannot reliably preserve delayed work, set batchSize: 1 and flushIntervalMs: 0 to send immediately.
Delivery is best effort. A rate-limited response (429) is retried with backoff — by default three retries per flush, each wait capped at five seconds, so a throttling window longer than that budget outlives it and the batch is dropped. Every other failure — a network error, a 5xx — drops its batch of up to 100 events immediately, with no retry. Drops are silent unless you pass a logger, and those warnings are sampled at 1% by default. Raise maxRetryAttempts and maxRetryDelayMs if your ingest key is throttled for long windows, maxQueuedEvents (1,000 events, oldest shed) if requests arrive faster than they flush, and errorSampleRate if you want to see drops as they happen.
Express and Node report the completed response status. A Next.js proxy cannot observe the final route status after pass-through, so the middleware helper omits it instead of guessing. When a Next.js proxy returns a response directly, use createDevTuneAiTraffic() and pass the known response.status; use withDevTuneAiTrafficRoute() where you own a Fetch-compatible route handler and want its exact status captured automatically.
Connect Google Analytics 4 and Search Console
Both connect by signing in with Google from Settings › Integrations: Google Analytics 4 under Website analytics, and Google Search Console under Search performance.- Google Analytics 4 — choose the property, the web stream, the tracked domain it maps to, and the key events you count as conversions. If you select no key events, sessions still sync but conversions stay empty rather than reading as zero.
- Google Search Console — choose the property. DevTune uses clicks and impressions to compare AI traffic against classic search, and to avoid recommending a rewrite to a page already earning Google traffic.
Install the Tracking Snippet
The snippet is a first-party record of AI referral page views. It does not see crawlers that skip JavaScript, so it is an addition to a crawler sensor rather than a substitute for one.Generate the Snippet
- Open your project
- Open Settings › Integrations
- Select Set up beside Tracking snippet under Website analytics
- Select Generate tracking snippet
Install the Snippet
Copy the snippet and add it to your website’s layout or template file so it loads on every page. The snippet should be placed before the closing</body> tag.
GitBook
If your documentation is hosted on GitBook, you can use the official DevTune integration instead of manually installing the snippet:- Go to your GitBook space settings
- Navigate to Integrations
- Find and install the DevTune integration from the marketplace
- Enter your project snippet key when prompted
Next.js
Add the snippet to your root layout file (app/layout.tsx or pages/_document.tsx):
WordPress
Add the snippet to your theme’sfooter.php file, or use a plugin such as “Insert Headers and Footers” to add it to the footer section.
Webflow / Squarespace
Go to Site Settings, then Custom Code, and paste the snippet in the Footer section.Plain HTML
Add a<script> tag before the closing </body> tag in your HTML template:
Enable or Disable Tracking
In Settings › Integrations, use the Snippet and middleware collection switch to enable or disable tracking at any time. The switch controls both the snippet and edge middleware; to stop only one, remove it from your site. When disabled, incoming events are rejected and no new data is recorded.Rate Limits
Each tracking snippet is rate-limited to 1,000 requests per minute. This is enough for most sites. If your site exceeds this limit, excess requests are silently dropped.Multi-Site Tracking
You can use the same tracking snippet across multiple websites or subdomains. Every site reports into the same project, and Insights › AI Traffic shows the total. To read one site on its own, pass its host in thedomain query parameter of the AI Traffic API — for example /api/v2/projects/{projectId}/traffic/summary?domain=docs.example.com. To read one path and its descendants, also pass pathPrefix — for example domain=example.com&pathPrefix=/blog.
Proxy Setup
You can proxy the tracking endpoint through your own domain when you need first-party routing for your deployment setup, or when ad blockers drop the third-party beacon. This applies to the snippet only — the crawler sensors do not use it. Teams should still honor consent requirements and local privacy rules before collecting traffic data.How It Works
Instead of the snippet sending data tohttps://devtune.ai/api/v1/llm-traffic/collect, you configure a reverse proxy, so requests go to a path on your own domain.
Vercel / Next.js
Add a rewrite tonext.config.ts:
Netlify
Add tonetlify.toml:
Cloudflare
Create a redirect rule or Worker that proxies/dt/* to https://devtune.ai/api/v1/llm-traffic/*.
nginx
Verifying Your Sensors
Each sensor confirms itself differently, and they do not all confirm at the same speed.
If the snippet reports no data, check:
- The snippet is present in your page source
- The snippet key matches the one shown in Settings › Integrations
- Ad blockers are not blocking the request
- Snippet and middleware collection is switched on in Settings › Integrations
- The tracking snippet sensor chip reports Active
A quiet site is not a broken one. A hostname that goes an hour without an AI
crawler hit still reports as Active — see how current each sensor
stays.
Next Steps
- Which Sensors Do I Need - Decide which sensors your project needs
- Reading AI Traffic - Explore your demand data
- AI Traffic Overview - Learn what gets tracked and why