Silktrace.
A complete product.
From interface to infrastructure.Site investigation, keyword research, technical audits and link prospecting in one workspace. I designed the product and public website, then built the application, persistence and controls around paid external data.
Bring the evidence together.
Make the next decision clearer.
Silktrace connects domain investigation, keyword research and technical site evidence in one workspace. I built the product, the systems behind its paid research requests and the website that introduces it.
- Designed & built by
- Alex Jardine
- The work
- Product, engineering & website
- Platform
- Web application
SEO research.
Site audits.
A place to act on both.
Silktrace is an SEO research and site-audit workspace for SEOs, growth marketers, consultants, agencies and site owners. It brings domain visibility, keyword demand, search results, technical crawl findings and link evidence together to help you decide what to fix, what to publish and where to investigate next.
Silktrace is built around the investigation between reports: move from a domain to a technical problem, from a keyword to the live SERP, or from a broken destination to the pages still linking to it, without rebuilding the context in another tool.
The principle: never invent certainty the source does not provide ↗See where a site already earns visibility.
Start with a domain and market. The overview brings its search footprint into focus before you commit to deeper research.
- Input
- A root domain or subdomain and a country/language market.
- Analysis
- Collect current organic metrics, top pages and available history. Compare domains sharing ranking keywords.
- Output
- Estimated organic traffic, known ranking keywords, a history chart, page contributions and competitor overlap.
- Decision
- Identify important pages, concentrated traffic and relevant competitors before choosing what to investigate.
From an issue count to the pages you can fix.
A technical audit starts with an explicit crawl request. The report leads with errors, warnings and notices, then lets you investigate the affected pages.
- Input
- A domain or saved project, page limit and crawl settings.
- Analysis
- Inspect crawlability, indexability, titles, descriptions, headings, canonicals, redirects and broken-link evidence. Hosted scans use DataForSEO OnPage.
- Output
- Grouped issues with affected counts, explanations, typical fixes, source and confidence; a separate Crawled URLs view.
- Decision
- Choose which faults deserve attention and find the URLs needed to make the repair.
Open the finding.
Issue details put the diagnosis beside the evidence. A duplicate-title report, for example, shows affected URLs, response status, title, indexability and canonical information.
Inspect individual pages.
Search Crawled URLs by URL or title. Filter by HTTP status, indexability or issue; sort by issue count, URL or response time. Compare title/H1, canonical and response evidence before changing a page.
Use the crawl again. Dead-link site mode reads broken internal and external destinations from an existing crawl, groups the evidence and shows occurrences. That turns technical evidence into a replace, redirect or remove decision without starting another crawl.
Understand the demand. Read the results.
Keyword research combines the size of the search opportunity with the pages already competing for it. A difficulty number alone does not tell you what to make.
- Input
- One keyword or a list of up to 50, with a regional language preset.
- Analysis
- Retrieve demand, CPC, difficulty, intent and trend; a single-keyword lookup also gathers SERP evidence.
- Output
- Metric rows; for a single query, ranking domains, titles, URLs, page-type labels, SERP features and People Also Ask questions.
- Decision
- Judge whether the demand, competition and preferred page format justify a brief or an update to existing content.
The saved Bishop lookup reports 90 monthly searches, difficulty 2 and commercial intent. But the results include hotels, booking sites and lodging roundups. The useful question is whether your page serves that lodging decision. An “easy” score is a research signal, not a promise of ranking. CPC adds paid-search context, not guaranteed organic revenue.
Expand the research: clusters, gaps and matrices
Keyword Cluster
- Input
- A seed keyword or topic.
- Analysis
- Find related terms and their demand, difficulty and intent.
- Output
- A sortable, filterable shortlist with row selection.
- Decision
- Choose which terms belong together in a brief; export selected or visible rows.
Competitor Gap
- Input
- A competitor domain or URL; optionally your client domain.
- Analysis
- Compare competitor ranking terms with available client coverage.
- Output
- Keywords, competitor positions and demand, with a gap indicator when comparison evidence supports it.
- Decision
- Find competitor topics worth validating for your own site.
Competitor Matrix
- Input
- Up to three competitors and an optional client domain.
- Analysis
- Combine their keyword evidence into one comparison.
- Output
- A keyword-by-domain ranking matrix with conditional client-gap filtering.
- Decision
- Spot shared competitive topics and compare where your client ranks or lacks coverage.
A plan with reasons.
Silktrace derives a 0–100 review score from demand, difficulty, intent and available client/competitor rankings. Each row carries a priority, suggested review action and explanation.
Keep editorial judgment.
“Review existing fit”, “Map to best page” and “Group with supporting content” guide the next decision. They do not automatically select a replacement page or create content. The plan can be exported separately.
Find the broken destination. See who still links to it.
Dead-link Opportunities connects topic discovery with destination checks, referring-link evidence and a possible contact route.
- Input
- A topic keyword, plus an optional replacement site or specific page URL.
- Analysis
- Find relevant candidate pages, check failure evidence, retrieve referring links and look for publisher contact evidence.
- Output
- A ranked table with dead destination, HTTP status, referring page, anchor/reason, rank/follow signals and contact route.
- Decision
- Check whether the source is relevant and the broken link is worth pursuing before approaching the publisher.
What makes a prospect interesting?
The priority score combines topic overlap, source rank, spam signals and follow status. The referring page and anchor give you the context to judge the link. Sourced contact emails show their kind and confidence; a suggested contact page is a fallback.
Suggestions you can inspect.
Supply your site or a specific replacement URL. When a saved crawl is available, Silktrace ranks eligible pages by term overlap in their title, H1 and URL path, and shows the matched terms. Suggestions must be HTTP 200 and indexable in that snapshot. A manual URL without supporting evidence stays unverified. You select the page after reviewing its content and intent; changing the replacement reads saved evidence without starting paid discovery.
A failed response may be temporary. The search covers a bounded set of candidates, and a contact route is not a verified inbox. The tool supports investigation; outreach remains your next step.
Inspect existing backlinks and AI search evidence
Backlink Research
- Input
- A domain.
- Analysis
- Request backlink totals and a bounded set of actual referring links.
- Output
- Up to 100 link rows with source/target URL, anchor, follow status, source/page rank, spam signals, dates and live/lost status.
- Decision
- Review who links to a site and assess the context behind its totals.
AI Search Evidence
- Input
- A domain and confirmed brand name, with optional aliases.
- Analysis
- Request separate brand-answer and domain-citation evidence.
- Output
- Questions, answer context, platform/model, cited and co-cited URLs, dates and available demand signals.
- Decision
- Inspect where and how the brand or site appears in returned answers. Coverage is bounded; ChatGPT evidence is US/English.
Save the domain. Return to the research.
Projects organize repeat site investigation. Saved Lookups preserve keyword research context. Each has a specific job.
- Input
- A domain to save, or a completed keyword lookup.
- Analysis
- Retain project context and stored research; separate reusable evidence from fresh provider work.
- Output
- Searchable projects and folders, saved reports, lookup mode/market/expiry and available historical snapshots.
- Decision
- Resume a review without rebuilding the context or automatically rerunning paid research.
Projects for sites.
Organize domains into folders, search and sort the list, archive old projects and reopen site reports. Project-scoped keyword research can live alongside the site investigation. Settings hold provider connection and research defaults in a private workspace.
Lookups for questions.
Overview, Cluster, Gap and Matrix searches create cache-backed lookup history. Reopen by query, mode and market while the cache remains reusable. An expiry date makes clear that a saved lookup is not a permanent promise of fresh data.
What can I compare over time?
Saved audit snapshots accept a label and note, then compare stored score, estimated visits and ranking-keyword counts. This comparison is a summary snapshot, not a diff of technical issues. Site scans and research run when requested; saving a project does not turn on scheduled rank tracking.
Never invent certainty.
A confirmed zero, an unrequested report and missing coverage lead to different decisions. The interface needs to make that distinction visible.
A reported value of zero is a result within the provider’s coverage. It is not proof that no demand exists anywhere.
No result has been measured yet. An unrequested backlink report cannot establish that a site has no links.
The technical report is still being prepared. The last completed snapshot remains distinguishable from new work.
Some evidence is absent or the request did not complete. Missing keyword metrics and unavailable client coverage should not be read as zero.
Volume, position, URL, HTTP status, backlink and anchor.
Issue grouping, page-type labels, opportunity score and suggested review action.
Inspect the source, choose the work, save the context or export the report.
Freshness, source and coverage keep those layers readable. The imported technical capture above explicitly shows stale and limited confidence. A bounded search returning no opportunities is not a complete census of the web.
Reports that become a brief or a fix list.
Choose the stored evidence you need. Explorer produces an Excel workbook; Keyword Research produces CSV files for the current research view or its plan.
| Handoff | What leaves Silktrace | Use it for |
|---|---|---|
| Site research · XLSX | Overview, Organic History, Top Pages, Organic Keywords and Competitors, according to selection. | Site review and comparison. |
| Technical audit · XLSX | Technical Summary, Technical Issues and Crawled URLs: severity, affected counts, typical fixes and page evidence. | An implementation queue. |
| SERP & links · XLSX | SERP Results and Features; Backlinks and Backlink Evidence; AI Search Evidence when available. | Reviewing search and link context. |
| Keyword research · CSV | Single-query metrics, SERP rows and PAA; bulk metrics or selected/visible cluster, gap and matrix rows. | An editorial shortlist or brief. |
| Keyword plan · CSV | Score, priority, suggested action, reason, available metrics and URL field. | A prioritized manual review. |
Workbooks include Export Info, selected report sheets, filterable headers and frozen first rows. Exports reuse stored evidence and do not start provider work. Missing reports remain unavailable; export limits apply. Dead-link prospecting currently has no dedicated export.
From a site question to a reviewable next step.
A hotel website provides a concrete example of how the tools fit together. You move between the research jobs; Silktrace keeps the evidence available for review.
- Open the saved domain.
Check which pages contribute estimated traffic and whether the keyword evidence covers the topic you care about.
- Inspect the technical report.
Open a duplicate-title issue and review affected URLs and canonicals before assigning a fix.
- Research the lodging question.
Open the saved Bishop lookup. Read volume, intent and the hotel/roundup results together.
- Choose the page work.
Use the plan’s reason to decide whether to improve an existing page, group supporting content or investigate further.
- Take the evidence into delivery.
Export technical evidence for implementation and keyword research or the plan for editorial review. Return through the project and Saved Lookups.
Silktrace adds the workspace, issue diagnosis, comparisons, review plans and report handoffs around DataForSEO evidence. Its value is a coherent way to investigate and decide, with fewer disconnected reports to reconstruct by hand.
Open SilktraceThe identity.
The interaction.
The implementation.
I designed and built Silktrace’s public website inside the application itself: a procedural silver identity, a reusable interaction system and a direct path into a private workspace.
Tailwind CSS 4 · Motion · WebGL
Next.js provides the route and layout boundaries. React composes the public experience; TypeScript types the application and auth contracts. Tailwind supports the product UI, while namespaced CSS builds the marketing system. Motion handles selected state transitions; native WebGL draws the identity.
Marketing components use JSX alongside the TypeScript application.
A quiet surface. A distinctive material.
The visual system pairs a pale green-grey canvas with near-black type and deep petrol actions. Silver gives the brand a material presence; the surrounding interface stays light enough for dense research screenshots to remain the evidence.
Inter carries both headings and body copy. Tight, balanced headlines contrast with open body leading, small technical labels and thin rules. Translucent navigation keeps the artwork visible; pill buttons give primary actions a softer edge.
- Canvas
- #F5F6F5
- Paper
- #FEFEFC
- Ink
- #14191D
- Petrol
- #032B3D
The layout rules behind the surface
A shared gutter uses max(28px, (100vw - 1248px) / 2). Section spacing scales from 76 to 124px; below 800px, the gutter becomes 26px and the shared section rhythm becomes 72px. Page-specific overrides refine that rhythm.
Paper, mist gradients and hairline borders separate chapters without surrounding every paragraph in a card. Lucide line icons identify capabilities. Screenshots distinguish illustrative sample figures from a real saved keyword report. The root loads Playfair too, but the public website explicitly uses Inter.
Source: frost.css · integration.css · app/layout.tsx · product-imagery.jsx
Silver, written in a fragment shader.
The hero is native WebGL on a canvas. Six vertices cover the drawing surface with two triangles; the fragment shader calculates the material at every pixel. There is no mesh, lighting rig or image sequence behind the silver loop.
Aspect-corrected coordinates become a radius and angle. Sine and cosine bend the ring, vary its width and turn it slowly. A narrow glint, fine ribs, an inner highlight and angular light falloff create the metallic reading. These are procedural shading terms, not physically based reflections.
float bend=.10*sin(a*3.+t)+.055*cos(a*5.-t*.6);
float ring=.58+bend;
float d=r-ring;
float twist=sin(a*2.-t*.8);
float width=.10+.075*(.5+.5*twist);
float body=exp(-pow(d/width,2.)*1.25);- CSS tokensBase, tone, highlight
- UniformsTime, size, pointer
- Fragment fieldBand, ribs, glint
- CanvasOne draw call
Pointer, color and animation lifecycle
getComputedStyle reads --shader-base, --shader-tone and --shader-light into RGB uniforms at initialization. Pointer coordinates are normalized against the canvas; each draw interpolates 3.5% toward the target and offsets the field by a small amount. Pointer input is ignored while paused or under reduced motion.
The animation accumulates time with a 50ms maximum step. ResizeObserver recalculates the drawing buffer, viewport and resolution uniform, then draws a frame without advancing time. IntersectionObserver, document visibility and user preferences gate the loop. Unmount cancels the frame, disconnects observers, removes listeners and deletes the buffer, program and both shaders.
The closing section reuses the renderer with a separate graphite mode: traveling waves move a stippled diagonal deposit, while a reading-zone mask thins the texture behind the heading. This mode has its own roughly 30fps draw throttle; the silver hero does not.
Source: shader.jsx · FrostMaterial in frost-interactive.jsx
Slow atmosphere. Quick decisions.
Continuous motion belongs to the material. Content arrives once, then stays still. Interface changes are shorter and smaller, so choosing a capability or opening an answer feels immediate.
- Section entrance 650ms / 22px
- IntersectionObserver triggers a Web Animations API fade and upward settle at 10% visibility. The easing, cubic-bezier(.16, 1, .3, 1), settles gently. The wrapper supports delay; the current pages do not orchestrate a stagger.
- Selected content 250ms / 6–8px
- Motion animates the product tab indicator and the changing content. Accordions animate height and opacity. Navigation uses a 300ms underline instead of a page-transition sequence.
- Action feedback 2px lift
- Buttons rise on hover and compress to .985 scale on press. Screenshot enlargement uses a native dialog, with focus returned to its trigger on close.
Keyboard and motion preferences
Product tabs have roving focus, Arrow keys, Home/End and a labeled panel. The mobile menu exposes expanded state, closes on navigation, and returns focus on Escape. Support search announces its result count; copy-email feedback uses a status region.
A skip link reaches the main landmark. Links navigate, buttons change state, decorative canvases are hidden from assistive technology, and focus-visible outlines use petrol. Reduced motion bypasses section reveals, removes CSS transitions and stops continuous shader drawing.
Source: Reveal · ProductDetails · Accordion · FrostHeader · CopyEmail
One application. Two visual environments.
The App Router’s public route group owns a header, main landmark and footer inside .fs-root. Private route layouts use HostedRouteBoundary; their product shell brings its own sidebar, denser layout and application tokens.
Marketing styles are ordinary CSS imported by the public layout, with selectors and tokens scoped under .fs-root. That matters during client navigation: even if those styles remain loaded, a marketing heading or button rule cannot match the private shell outside that root.
/ · /product · /access · /engineering
/support · /security · /privacy · /terms · /login
/explore/* · /projects/*
/research · /settings · /site-scan
Reusable behavior, page-specific composition
frost-pages.jsx composes each public page. FrostButton, TextLink and Eyebrow establish recurring controls and type roles; Reveal handles entrances; FrostMaterial wraps the shader and pause preference. ProductImagery owns screenshot selection, provenance and enlargement. ProductDetails and SupportFinder keep their own local interaction state.
The shared root supplies fonts and global styles. Marketing namespaces --fs-* and --premium-* avoid replacing product variables; local reset and heading rules account for the shared base. This is selector and token isolation, not CSS Modules, Shadow DOM or a guarantee of separate CSS bundles.
The public login reuses the application’s typed SocialSignIn component, passed into its marketing presentation. Server page composition stays separate from client components that need state, browser observers or canvas access.
Source: app/(public)/layout.jsx · integration.css · hosted-route-boundary.tsx · app-shell.tsx
Recompose the scene. Bound the work.
Desktop places the silver beside a large, left-aligned heading. On mobile, the artwork moves above the copy and a vertical veil protects the reading area. It is a different composition using the same shader.


- Desktop
- Fluid 62–102px hero type, a centered 1248px content measure and paired text/media columns.
- Tablet / ≤1000px
- Gaps and panel padding tighten before the main 800px layout change. Columns are retained where space permits.
- Mobile / ≤800px
- 40–66px hero type, stacked sections, a disclosure menu with 60px rows and 48px hero buttons. Sign-in omits its artwork. Gallery previews crop intentionally; enlargement restores the full report, while the saved-keyword feature image scales without cropping.
What the rendering controls actually do
The existing shader caps the buffer ratio at min(devicePixelRatio, 1, 1100 / width), limiting its width to 1100 pixels. It requests a low-power WebGL context without antialiasing. This bounds fragment work, with an explicit sharpness tradeoff on high-density screens.
The loop stops offscreen, in a hidden tab, when paused or with reduced motion. Resize still draws a static frame. Each material restores its pause preference from sessionStorage on mount. Context loss switches to a static fallback; automatic context restoration is not implemented.
Product captures use next/image with intrinsic dimensions and responsive sizes. Next/font loads Inter and Playfair with swap behavior; marketing uses Inter. Browser APIs live in client effects. The public layout is force-dynamic: it is not a fully static export. The inspected marketing components do not use explicit dynamic imports or a lazy-mounted shader.
Source: integration.css · frost-pages.css · shader.jsx · product-imagery.jsx · app/layout.tsx
See the work. Then open a workspace.
The product gallery makes the interface inspectable before an account is required. Overview and page captures label their illustrative data; the keyword capture is a real saved report. Switching or enlarging an image requests an image asset, not a research endpoint.
The standalone sample was retired in 0.1.47. Its URL permanently redirects to the product page. The current public demonstration is screenshot-based, with no editable demo workspace or paid provider action.
- Public CTAOpen your workspace → /login
- Google through Neon AuthPending state, disabled button, recoverable error
- /auth/continueIncomplete setup → onboarding
Otherwise → Domain Explorer
The real sign-in behavior
SocialSignIn calls Neon Auth’s redirect-based social sign-in with Google. Both normal and new-user callbacks point to /auth/continue; errors return to /login?error=oauth. The client uses a 20-second request timeout and an alert for failures.
The server redirects already signed-in visitors away from /login. The continuation route reads onboarding progress and chooses /onboarding or /explore/overview. The public header itself always says Sign in; it does not swap navigation based on auth state. An arbitrary originally requested destination is not preserved by this flow.
Source: next.config.ts · product-imagery.jsx · social-sign-in.tsx · login/page.tsx · auth/continue/page.tsx
Explore the product websiteBased on the 0.1.48 repository and a live public-route review on 8 September 2026.
React 19. Next.js 16.
PostgreSQL. Drizzle.
Neon Auth. DataForSEO.
A custom orchestration layer connects them.
I built the systems that turn external evidence into a dependable hosted product: bounded provider pipelines, inspectable ranking, private workspaces and durable controls around paid research.
The application boundary
One application, with clear ownership of each responsibility.
Research controls · report views · Recharts
Neon session → WorkspaceScope · bounded input
Request identity → cache → estimate → confirmation → reservation
hosted-dead-link-servicehosted-keyword-research-servicehosted-technical-audit-serviceCredentials, quota ledger, task state, cached reports and crawl snapshots
Fixed provider origins, server credentials, timeouts and response mapping
| DataForSEO provides | Silktrace adds |
|---|---|
| Content search results | A bounded topic-to-candidate workflow, URL validation and candidate deduplication. |
| Page status and link evidence | Failure qualification, conditional backlink enrichment and separate repair/prospecting modes. |
| Backlink records and parsed content | Publisher contact routing, provenance and deterministic opportunity ranking. |
| Keyword metrics and SERPs | Market-aware research modes, a normalized keyword model and planning rules. |
| Individual API responses and costs | Signed approvals, quota reservations, submission replay, workspace isolation and reusable reports. |
Deep implementation detailsRoutes, service boundaries and execution model+
Research controls · report views · Recharts
Neon session → WorkspaceScope · bounded input
Request identity → cache → estimate → confirmation → reservation
hosted-dead-link-servicehosted-keyword-research-servicehosted-technical-audit-serviceCredentials, quota ledger, task state, cached reports and crawl snapshots
Fixed provider origins, server credentials, timeouts and response mapping
The React client sends research intent to routes such as /api/research and /api/explore/dead-links. Next.js resolves the authenticated owner, validates the request and calls a service with a server-created WorkspaceScope. The service selects reusable evidence or coordinates new paid work through a DataForSEO adapter.
Drizzle maps records in the PostgreSQL silktrace schema. Repository methods receive that same scope for reads, writes and task transitions. Adapters return internal results; services persist them and assemble view models for the interface. Most live research runs inline during the route request. The technical crawl has a separately persisted provider task that subsequent requests can advance. The queue module currently reports inline mode; this architecture does not depend on an active background worker.
| DataForSEO provides | Silktrace adds |
|---|---|
| Content search results | A bounded topic-to-candidate workflow, URL validation and candidate deduplication. |
| Page status and link evidence | Failure qualification, conditional backlink enrichment and separate repair/prospecting modes. |
| Backlink records and parsed content | Publisher contact routing, provenance and deterministic opportunity ranking. |
| Keyword metrics and SERPs | Market-aware research modes, a normalized keyword model and planning rules. |
| Individual API responses and costs | Signed approvals, quota reservations, submission replay, workspace isolation and reusable reports. |
The custom work lives between provider calls: deciding which result becomes the next request, bounding that fan-out, reconciling inconsistent fields and retaining enough evidence to review the output. This chapter describes the hosted implementation reviewed on 8 September 2026. Legacy local storage and sample data paths are separate.
Implementationsrc/app/api/research/route.ts src/server/services/hosted-*-service.ts src/server/sync/queue.ts
Controlling paid work
Authorization, concurrency and accounting belong on the server.
- Client intentResearch parameters + submission key
- Server estimateResolve cost and confirmation policy
- Signed approvalBind workspace, request, amount and key
- Transactional reservationClaim attempt + quota under a lock
- Provider executionRun the bounded dependency graph
- SettlementClose task + reconcile quota and cost
- Usage receiptRetain estimated and returned spend
Invalid or expired approval Return a new quote before execution.
Same submission replayed Return the existing attempt or completed evidence.
Stale or uncertain execution Reconcile quota and record unknown spend; do not blindly retry.
Reservations own quota, not provider-account funds. An estimate is not a hard spending cap; returned cost can remain unknown.
An HMAC-signed quote binds approval to a workspace, request, cost and idempotency key. A PostgreSQL transaction then claims the attempt and its quota slots under a workspace advisory lock. Failed checks roll back before provider execution.
Replaying the same submission cannot start a second pipeline. Settlement closes the task, releases concurrency and records usage. Dead-link enrichment runs inside one durable attempt; it does not checkpoint every individual HTTP call.
Deep implementation detailsQuote payload, quota rules, state machine and failure semantics+
Estimate → bind → confirm
evaluateProviderCostConfirmation() requires confirmation for estimates at or above US$0.05. Dead-link research uses a fixed US$0.25 estimate, rather than dynamically pricing its eventual fan-out. Labs estimates use 0.01 + max(1, itemCount) × 0.0001; standard SERP estimates use 0.002 × max(1, ceil(depth / 10)). These are constants in the application, not a current provider price quote or a hard spending cap.
stc1.<base64url payload>.<HMAC-SHA256 signature>
{
format, version, keyVersion,
workspaceHash, // SHA-256 of server-resolved workspace
operation, requestHash,
idempotencyHash, // binds approval to this submission
costMicros, // server-computed amount
issuedAt, expiresAt, // five-minute validity
nonce // random value
}A server-only, versioned signing key authenticates the payload with a domain-separated HMAC. Verification checks strict schema, canonical encoding, constant-time signature equality, expiry and each binding against freshly computed server context. Modifying the amount, request, credential-dependent hash or submission key invalidates approval. A client boolean alone cannot authorize hosted work.
The initial challenge returns HTTP 409 without reserving a task or calling DataForSEO. Confirmation resubmits the token and the same idempotency key. Credential rotation or changed parameters require a new valid quote; an invalid/expired token receives a replacement challenge.
What the reservation actually reserves
reserveProviderTask() executes a PostgreSQL transaction behind a workspace advisory lock. It verifies the current credential version, claims idempotency_records, increments quota reservations, inserts a provider_tasks row and records each quota slot in provider_task_quota_reservations. A failed check rolls the transaction back.
| Resource | Reservation behavior |
|---|---|
| provider_action_daily · 25 | One slot per logical provider action, not per nested HTTP call. |
| expensive_concurrency · 1 | One active expensive task per workspace. Always released at settlement. |
| technical_audit_daily · 3 | An additional daily slot for requests with the technical quota class. |
The atomic quota update succeeds only when reserved + used + 1 ≤ limit. These records reserve permission to execute within usage limits. They do not hold money or reserve funds in a DataForSEO account. Settlement subtracts reserved slots, consumes daily slots if work may have reached the provider, releases concurrency, records cost and closes the idempotency record in one transaction.
failedcancelledexpiredsubmitted can go directly to collecting. polling and collecting permit self-transitions. Terminal records cannot be settled twice. There are no database states named planned, confirmed, running or partial.For dead-link research, the task remains submitting while the complete live pipeline runs. Only afterward does it move through submitted and collecting, save the result and settle. Its individual live calls are not persisted as separate provider tasks. The asynchronous OnPage crawl does persist its remote task ID for later polling and collection.
| Situation | Actual behavior |
|---|---|
| Same workspace + operation + idempotency key | The unique database identity finds the existing attempt. Its request hash must match; different parameters conflict. |
| Existing attempt still in progress | Dead-link service returns HTTP 202 with the existing task ID. No second pipeline is started. |
| Existing attempt completed | Returns the cached result. If the completed task has lost its evidence, it reports a reconciliation error instead of executing again. |
| Existing attempt failed, cancelled or expired | Returns HTTP 409 requiring a new request. There is no restart from the last dead-link enrichment stage. |
| Same search, different submission keys, cold cache | Not universally deduplicated by request hash. The concurrency quota blocks overlapping expensive work; it does not attach both submissions to one task. |
| Attempt becomes stale | Reconciliation uses a 15-minute stale threshold against next-attempt/update time. Never-submitted reservations release daily slots. Possibly submitted work consumes daily quota, releases concurrency and records unknown spend. No blind provider retry. |
Accounting for evidence and uncertainty
extractDataForSeoCost() prefers a numeric response-level cost; otherwise it sums task costs. That avoids counting both the envelope and its tasks. provider_usage_receipts stores the task, endpoint identity, query hash, mode, estimated/actual USD and call time when a reservation is consumed.
The dead-link adapter sums costs across its calls and returns null if any call’s cost is unknown. Keyword and crawl adapters can sum the known subset instead, so a non-null total is not universally proof of a complete bill. attempt_count increments on task transitions; it is not a network retry counter. These distinctions keep request state, report coverage and accounting uncertainty separate.
Each result earns the next request
The dependency graph turns a topic into evidence worth reviewing.
- DiscoverContent Analysis Search · up to 20 candidates
- QualifyNormalize → deduplicate → cap each host at 5
- VerifyInstant Pages · retain first 5 HTTP 404/410 destinations
- Find publishersBacklinks Live · up to 20 links per destination
- Find a contactContent Parsing · first 10 distinct source domains
- Rank & retainScore → sort → keep up to 50 rows → persist
Up to 17 conditional HTTP requests. One discovery, one verification batch, up to five backlink lookups and ten contact parses.
Only confirmed HTTP 404/410 destinations reach backlink enrichment. Referring pages supply link context and possible publisher contacts. Every stage is bounded, and missing checks remain visible in coverage warnings. Contact evidence identifies a route to investigate, not a verified inbox.
Deep implementation detailsProvider payloads, mapping fallbacks, limits and timeouts+
Keyword Mode starts with a topic, an optional replacement site and an optional specific replacement URL. DeadLinkOpportunitiesClient posts to /api/explore/dead-links. It creates and retains an idempotency key through submission and confirmation; changed input invalidates that key. An in-flight guard prevents overlapping client submissions. There is no language/location selector or user-configurable research budget in this request.
POST /api/explore/dead-links
Idempotency-Key: <one key retained through confirmation>
{
"mode": "keyword",
"keyword": "sustainable building materials",
"replacementTarget": "example.com",
"replacementUrl": "https://example.com/materials-guide"
}
// The confirmation submission adds:
// confirmationToken: <signed server quote>
// confirmCost: trueThe route rejects unknown fields and bodies over 8 KiB. normalizeDeadLinkKeywordInput() normalizes Unicode and whitespace and accepts a 2–200 character topic. Replacement URLs must be public HTTP(S), without embedded credentials; fragments are removed. Credentials, workspace identity, endpoint choices and execution limits come from the server.
- DiscoverContent Analysis Search · up to 20 candidates
- QualifyNormalize → deduplicate → cap each host at 5
- VerifyInstant Pages · retain first 5 HTTP 404/410 destinations
- Find publishersBacklinks Live · up to 20 links per destination
- Find a contactContent Parsing · first 10 distinct source domains
- Rank & retainScore → sort → keep up to 50 rows → persist
Discover candidate resources
POST /v3/content_analysis/search/live {
"keyword": "sustainable building materials",
"search_mode": "one_per_domain",
"limit": 20,
"rank_scale": "one_hundred",
"order_by": ["domain_rank,desc"]
}The adapter reads tasks[0].result[0].items. It takes url, falling back to page_url, and a title from content_info.title, main_title or item.title. Domain rank orders discovery; it is not the final opportunity score.
candidates() normalizes public URLs, removes fragments, drops exact duplicates and retains at most five URLs per normalized host and twenty overall. Lowercasing and removing www. unify host identity. Query parameters and HTTP/HTTPS variants are not fully canonicalized. No language, location or extra content filters are sent by this adapter.
Check whether a candidate is unavailable
POST /v3/on_page/instant_pages {
"url": "https://publisher.example/old-resource",
"store_raw_html": false,
"load_resources": false,
"enable_javascript": false,
"return_despite_timeout": true
}verification() requires task status 20000, associates each task with task.data.url or its input index, and checks that the URL belongs to the submitted candidate set. It reads the first page item and qualifies only HTTP 404 or 410. Access errors, rate limits, server errors and unknown status are not treated as confirmed dead destinations.
Only the first five confirmed destinations continue, in discovery order. The report counts successful checks and all confirmed dead pages separately; missing checks and enrichment truncation produce warnings. These limits also match the endpoint’s documented batch allowance of twenty URLs and five identical domains. Instant Pages contract ↗
Find pages still linking to each failed destination
POST /v3/backlinks/backlinks/live {
"target": "https://publisher.example/old-resource",
"limit": 20,
"exclude_internal_backlinks": true,
"backlinks_status_type": "live",
"rank_scale": "one_hundred",
"order_by": ["rank,desc"]
}backlinkRows() maps the first result’s items into referring-page evidence: source URL/domain, anchor, follow status, source rank, page rank and spam score. It rejects invalid source URLs, lost links and records whose normalized url_to does not equal the verified destination. Alternative provider names such as url_from/source_url and rank/domain_from_rank map to one internal field.
A Map deduplicates the destination/source pair before contact enrichment. Liveness still comes from the provider index; the referring page’s link is not independently rechecked. Spam lowers the later score without a hard rejection threshold. Failed enrichment adds a warning while other destinations can continue.
Extract an attributable contact route
POST /v3/on_page/content_parsing/live [{ "url": "https://source.example/article-linking-to-resource" }]The service takes the first referring page for each distinct source domain, then parses at most ten domains sequentially. Selection happens before opportunity sorting. It recursively reads returned strings to depth twelve with an output-count guard around 2,000 strings, extracts emails with a regular expression and deduplicates lowercase addresses.
Emails must belong to the source domain or a parent domain. Editorial/outreach role addresses rank first, named addresses second and generic inboxes third. The chosen email has medium confidence and records its source page. This is domain-associated evidence, not mailbox verification. No email yields an unknown contact kind with low confidence; a failed or unrequested parse leaves confidence unknown. Both use the publisher homepage as a manual-research fallback. Contact evidence is shared by rows from the same domain.
Implementationsrc/app/api/explore/dead-links/route.ts src/server/services/hosted-dead-link-service.ts src/server/providers/dataforseo/hosted-dead-link-prospecting.ts
The score is inspectable
Deterministic heuristics make each ranking explainable.
- R · Relevance
- round(100 × matched terms / unique terms)
- A · Authority
- sourceRank ?? pageRank ?? 0
- Q · Quality
- max(0, 100 − (spamScore ?? 35))
- F · Follow signal
- 10 for followed links; 3 otherwise
A score of 91, explained.
Every topic term matches, source rank is 72, spam score is 4 and the link is followed.
45 + 21.6 + 14.4 + 10
Topical relevance measures term overlap in the destination title, URL, anchor and referring URL. Source authority, spam and follow evidence complete the score. It prioritizes review; it does not predict outreach success.
Replacement fit is a separate score.
Saved crawl pages are ranked by exact term overlap in titles, H1s and URL paths. Eligible suggestions must be HTTP 200 and indexable in the snapshot. The interface exposes matched terms and leaves selection to the user. Unsupported manual URLs stay unverified. Changing the replacement reads saved evidence without repeating paid discovery.
Deep implementation detailsTokenization, defaults and replacement-matching algorithm+
| Provider-derived evidence | Silktrace-derived values |
|---|---|
| Destination title/status; referring URL and anchor | Normalized URL and domain identity; stable row ID |
| Source rank, page rank, spam score, follow flag | Topical relevance and weighted opportunity score |
| Parsed page strings containing contact evidence | Email classification, domain association and contact confidence |
| Saved crawl URL, title, H1, HTTP status and indexability | Replacement candidates, matched terms and overlap percentage in a separate storage-only service |
relevance() lowercases the topic, splits on non-ASCII letters/numbers and keeps unique terms longer than two characters. It checks whether each term appears as a substring in the destination title, destination URL, anchor or source URL. Relevance is the rounded percentage of matched terms, or zero if there are no qualifying terms. Body content, embeddings and an LLM are not used.
- R · Relevance
- round(100 × matched terms / unique terms)
- A · Authority
- sourceRank ?? pageRank ?? 0
- Q · Quality
- max(0, 100 − (spamScore ?? 35))
- F · Follow signal
- 10 for followed links; 3 otherwise
A score of 91, explained.
Every topic term matches, source rank is 72, spam score is 4 and the link is followed.
45 + 21.6 + 14.4 + 10
Numeric rank/spam evidence is bounded to 0–100. Unknown authority contributes zero; unknown spam assumes 35. Deduplicated rows sort by score descending and stop at fifty. There is no minimum-score cutoff or high/medium/low category in this dead-link path. A 24-character SHA-256 prefix supplies the stable destination/source row ID.
Replacement matching is a separate evidence join
After the research service returns, the route calls enrichDeadLinkReplacements() for both GET and POST. It resolves a replacement domain from the requested site, Site Mode target or manual URL hostname, then loads up to 1,000 pages from that domain’s latest workspace-owned technical snapshot. A sandbox snapshot is excluded unless the opportunity report is also in sandbox mode. It reads existing storage; changing a replacement does not initiate paid provider work.
- Wanted termsTopic + dead title + anchor + dead URL path
- Eligible pagesSame host · HTTP 200 · indexable · not the dead URL
- Available termsCandidate title + H1 + URL path
- Rank candidatesExact term overlap → top three → user selection
wanted = uniqueTokens(keyword + deadTitle + anchorText + deadPath)
available = uniqueTokens(page.title + page.h1 + page.path)
matchedTerms = wanted.filter(term => available.has(term))
score = round(100 * matchedTerms.length / wanted.length)
// Reject candidates with no matched terms.
// Sort by score descending, then URL for deterministic ties.
// Retain the top three; selection remains with the user.The tokenizer normalizes Unicode, lowercases, splits on non-letter/non-number characters and removes short terms plus a small stop-word list. Eligible pages must be on the replacement host, HTTP 200 and indexable in the saved crawl. Duplicate URLs and the dead destination itself are excluded. Candidate titles, H1s and paths contribute exact tokens; body content, semantic embeddings and LLM judgments do not.
For four wanted terms, matching three gives 75% overlap. That number is an inspectable lexical fit, not a probability that the replacement satisfies the same intent. Each suggestion retains matchedTerms and evidence:"saved-crawl"; the response also exposes crawl date/context. There is no high/medium replacement-confidence classifier.
A manual URL is placed first unless it is the dead destination itself. If it is already an eligible scored candidate, its saved evidence is retained; otherwise it is labeled manual-unverified with a null score. Up to three other suggestions can follow. With no matches and no manual URL, the list stays empty and the UI offers a crawl refresh or a supplied URL. Suggested pages are not selected automatically. The provider cache stores base opportunities; replacement suggestions are recomputed from stored crawl evidence on response, so a new page choice does not invalidate paid discovery.
Implementationhosted-dead-link-prospecting.ts · relevance(), contact(), executeHostedDeadLinkKeywordProspecting() src/server/services/dead-link-replacement-service.ts · enrichDeadLinkReplacements(), rankDeadLinkReplacements() src/lib/contracts/page-vms.ts · DeadLinkReplacementVM
Different questions, different pipelines
One normalized keyword model carries the results back to the interface.
- OverviewKeyword metrics; a single lookup also gets a SERP
- ClusterMerge expansions and suggestions, then enrich terms
- GapJoin competitor rankings with client overlap
- MatrixBuild a keyword union across competing domains
Each mode has its own provider plan and mapper. The browser receives the same application concepts: demand, difficulty, intent, rankings and evidence gaps. A requested keyword missing from a bulk response gets a no-data row.
S = clamp(round(V + D + I + G + C), 0, 100)
- V · Demand
- Log-scaled volume · up to 35
- D · Difficulty
- Lower difficulty · up to 25
- I · Intent
- Search intent · up to 20
- G · Client ranking
- Available position · up to 15
- C · Competitor ranking
- Competitive position · up to 10
Missing client ranking is missing returned evidence, not proof of absence. CPC is shown as context and does not enter the score.
Deep implementation detailsEndpoint plans, exact scoring rules and normalization+
/api/research resolves settings and optional project ownership before execution. Region presets become concrete provider codes: US/English is 2840/en, Canada/English is 2124/en, and Canada/French is 2124/fr. Missing/invalid selection falls back through workspace defaults. Overview splits comma, newline or semicolon input, collapses whitespace and deduplicates case-insensitively, up to fifty keywords.
| Mode | Provider dependency and parameters |
|---|---|
| Single overview | dataforseo_labs/google/keyword_overview/live: keywords[], location/language and include_clickstream_data:false. Then serp/google/organic/live/advanced: keyword, same market, desktop, depth 10, async AI overview disabled. |
| Bulk overview | The same Labs overview endpoint with up to 50 keywords. No separate SERP request per keyword. |
| Cluster | keywords_data/google_ads/keywords_for_keywords/live with seed and limit 200, plus dataforseo_labs/google/keyword_suggestions/live with limit 200. Merge unique candidate terms; enrich the first 100 through keyword overview when candidates exist. |
| Gap | dataforseo_labs/google/ranked_keywords/live for the competitor, organic rows, limit 200; then dataforseo_labs/google/domain_intersection/live for competitor/client, intersections enabled, limit 1000. Join overlap evidence by normalized keyword. |
| Matrix | Sequential ranked_keywords/live requests, limit 1000 per unique competitor/client domain. Build one keyword union with position and URL per domain. |
HostedResearchRawPayload is a discriminated union: single overview carries overview + SERP; cluster carries expansion + suggestions + enrichment; gap carries ranked + intersection; matrix carries results by domain. mapHostedResearchPayload() selects the matching mapper rather than asking the UI to interpret different provider structures.
Input: { tab: "overview", query: "where to stay in Bishop", region: "us-en" }
→ contract: keywords[], locationCode: 2840, languageCode: "en"
→ workspace cache lookup; on miss, estimate/reserve/execute
→ keyword_info.search_volume / cpc + keyword_properties + intent
→ normalized metrics: monthlyVolume: 90, cpc: 2.96,
keywordDifficulty: 2, intent: "commercial"
→ buildOpportunity(): score 73, priority "high"
→ mapped JSONB result + best-effort saved lookup
→ KeywordResearchPageVM → metric views, plan rows and trend chartsThe mapper tolerates metrics nested under keyword_info, root fields or keyword_data.keyword_info. Volume/difficulty become bounded numeric values or null; unknown intent is explicit. Cluster merging fills missing values across sources by normalized keyword. Monthly trends are sorted by year/month and retain the latest twelve entries. A requested bulk keyword absent from the response gets an explicit no-data row instead of disappearing.
Keyword planning uses a different score
buildOpportunity() adds five components, rounds and clamps to 0–100. CPC is displayed but does not enter this formula.
| Component | Exact rule |
|---|---|
| V · Demand | min(35, log10((monthlyVolume ?? 0) + 1) × 8). Log scaling limits the advantage of very large search terms. |
| D · Difficulty | Missing difficulty: 12. Otherwise max(0, (100 − difficulty) × 0.25). |
| I · Intent | Transactional 20; commercial 18; informational 11; navigational 4; unknown 7. |
| G · Client ranking | Missing position 15; positions 4–20: 12; beyond 20: 8; top three: 3. Missing here means absent from returned evidence, not proven absent from all search results. |
| C · Competitor ranking | Missing position 0; top three 10; positions 4–10: 8; positions 11–20: 4; beyond 20: 1. |
| Priority | High at 70 or above; medium at 45–69; low below 45. |
For the illustrated commercial keyword, volume 90 contributes 15.67, difficulty 2 contributes 24.5, intent contributes 18, missing client position contributes 15 and missing competitor position contributes zero. The total rounds to 73.
Recommended actions are a separate ordered rule set: difficulty ≥75 with volume <500 gives “Low priority”; a client at 4–20 gives “Review existing fit”; top-three client rankings give “Validate manually”; informational intent with missing difficulty or difficulty ≤55 gives “Group with supporting content”; missing difficulty or difficulty ≤65 gives “Map to best page”; otherwise, “Validate manually.” Thus a score category and an action can differ. Existing ranking URLs can be carried forward, but no semantic page matcher is invoked.
The mapped result is cached for thirty days under its market and credential identity. A single overview’s built-in estimate is US$0.0121, below the confirmation threshold; a cluster’s is US$0.125, above it. Both still use durable reservations and usage recording when they execute paid work.
Implementationsrc/server/providers/dataforseo/hosted-keyword-research.ts src/server/services/research-service.ts · mapHostedResearchPayload(), buildOpportunity() src/lib/keyword-research.ts src/components/research/keyword-research-client.tsx
Ownership reaches the database
Private workspaces are enforced beyond the interface.
Workspace-qualified foreign keys. A task cannot point to another workspace’s credential, and research records retain their owner relationships.
Neon Auth establishes the session. The server creates the workspace scope and carries it through services and repository queries. Composite relational constraints protect the links between credentials, projects, tasks and crawl evidence.
Credential changes advance a version and invalidate associated cache entries. Execution checks that version again so an approved request cannot silently use a different account configuration.
Deep implementation detailsAuthentication, schema relationships and credential lifecycle+
Neon Auth handles sign-in and session cookies. On the server, getCurrentUser() calls auth.getSession() and requires a user ID, session ID and valid session creation time. getHostedWorkspaceSession() calls bootstrapWorkspace(), then creates WorkspaceScope.fromAuthenticatedContext() from that user, owner membership, workspace, session and a generated request identity. The browser does not choose the workspace ID for a research operation.
After sign-in, /auth/continue resolves the workspace and reads its settings. Incomplete setup goes to onboarding; completed setup goes to Domain Explorer. Missing sessions return to login, while deleted or suspended accounts receive their own access outcome. Paid API handlers resolve authorization independently rather than trusting that a page was previously opened.
Bootstrap checks existing active membership and workspace status. A transaction takes a user-specific advisory lock and rechecks before creating the workspace, owner membership and settings. This avoids two simultaneous first requests creating two owner workspaces. WorkspaceScope has a private constructor and validates owner/session context; repository predicates still provide the actual data-access boundary.
Research execution
provider_credentialsEncrypted secret · version · modeidempotency_recordsUnique workspace + operation + keyprovider_tasksCredential + idempotency references
Status · remote task ID · costsprovider_task_quota_reservationsTask ↔ quota_windows · reserved/used/releasedprovider_usage_receiptsTask-linked usage and provider cost
Reusable results
provider_cacheCredential/version · request hash
JSONB result · expiration · logical bytessaved_research_lookupsCache reference · label · search parameters
Optional project referenceprojectsWorkspace-owned domain/project context
Technical evidence
technical_audit_runsTarget · crawl configuration · task referencetechnical_audit_snapshotsOne saved snapshot per runtechnical_audit_pagesURL · metadata · status · page metricstechnical_audit_issuesCode · severity · recommendationtechnical_issue_occurrencesIssue ↔ page · JSONB evidence
Workspace-qualified foreign keys connect records within the same owner boundary. Provider cache rows and technical snapshots are different persistence systems.
Queries filter on scope.workspaceId. Composite foreign keys such as (workspace_id, credential_id) prevent a child record from referring to another workspace’s credential; equivalent constraints protect projects, tasks and crawl evidence. This is application-enforced tenant scoping backed by relational constraints, not a claim that database row-level security is enabled.
Credentials are versioned execution dependencies
provider_credentials has one record per workspace/provider. Secrets use AES-256-GCM encryption with workspace/provider/field context authenticated alongside the ciphertext. The server decrypts them only for provider access. The row also stores live/sandbox mode, verification status, encryption key metadata and a monotonically changing credential version.
Replacement, mode changes or disconnection advance the credential version and clear its provider cache. Execution rechecks the expected version before using the secret. This prevents an approved request from silently running against a different account configuration and prevents evidence from the old credential context being reused as if it belonged to the new one.
Implementationsrc/server/auth/session.ts src/server/auth/hosted-workspace-session.ts WorkspaceScope.fromAuthenticatedContext() postgres-hosted-repository.ts · bootstrapWorkspace() src/server/db/schema/hosted.ts
Reuse requires an identity match
Fresh evidence is reusable only inside the right ownership and request context.
The cache checks those identities and expiry before reusing a report. Market changes, rotated credentials and changed request contracts lead to a miss. Saved keyword lookups point to reusable cached evidence; they are not permanent archives.
Technical snapshots have separate persistence. They support issue review, Site Mode and replacement suggestions without launching a new crawl. Capture dates keep that distinction visible.
Deep implementation detailsCache keys, freshness, storage bounds and saved lookups+
(workspace_id,
credential_id,
credential_version,
endpoint_key,
request_hash,
data_mode)
AND expires_at > now()| Workflow | Hashed input and freshness |
|---|---|
| Dead-link Keyword Mode | SHA-256 of JSON containing endpoint identity dataforseo.dead-link-prospecting.v2, credential ID/version, keyword and fixed 20/5/20/10 limits. Replacement site/URL are excluded because their enrichment reads storage only. TTL: 7 days. |
| Keyword research | Credential context plus project ID and versioned contract: tab, query, keyword list, location/language codes and matrix domains. TTL: 30 days. gapOnly is presentation filtering, not provider identity. |
| Organic / backlink research | Their own endpoint and normalized request identities. Organic bundle TTL: 24 hours; backlink evidence TTL: 7 days. |
Workspace and live/sandbox mode are separate key components even when they are not inside the hash. A global cache could expose one owner’s research or incorrectly treat another owner’s credential-bound request as already paid. Credential versions also prevent reusing old provider-account context after rotation.
Expired rows, different markets, changed inputs, a new endpoint version or rotated credentials produce a miss. Normalization is deliberately bounded: dead-link keyword case is preserved for identity, so two case variants can still miss the cache even though the relevance calculation lowercases both.
Saved evidence has an explicit lifetime
The cache stores mapped result JSONB, not a raw provider envelope or a normalized opportunity table. A dead-link payload contains scored base rows, coverage counters, warnings and capture time. A keyword payload contains the mapped research result. Cache writes enforce a 5 MiB entry limit and a 50 MiB workspace storage budget, with expiration cleanup and logical-byte accounting.
saved_research_lookups adds a label and search parameters pointing at a cache row. Reopening a reusable lookup can render stored evidence without provider work. That pointer is not an immutable archive: deleting its cache row cascades to the lookup. Keyword history creation is best-effort after cache persistence, so a usable result can exist without a recent-history entry. Dead-link cache writes do not create this saved-lookup record.
Technical reports instead read persisted crawl snapshots. Site Mode derives broken links from the latest snapshot without refreshing it. Its displayed seven-day timestamp does not enforce rejection of an older snapshot; users are viewing captured evidence, not a fresh status check.
Implementationpostgres-hosted-provider-store.ts · getProviderCache(), saveProviderCache() hosted-keyword-research-service.ts · getHostedResearchRequestContract() src/server/db/schema/hosted.ts · providerCache, savedResearchLookups
One request, end to end
A cold dead-link search connects the architecture to the user’s next decision.
- Ask. Send a topic and submission key. The server resolves ownership, validates input and checks reusable evidence.
- Approve and reserve. On a miss, confirm the signed estimate. Claim the durable attempt and quota before spending.
- Investigate. Discover candidates, confirm failed destinations, find referring pages and collect contact evidence.
- Save and settle. Rank the rows, persist coverage and warnings, settle the task and record usage.
- Choose. Join saved-crawl replacement suggestions. Return one report with sources, scores and choices the user can inspect.
Deep implementation detailsThe full request walkthrough+
- Construct and validate. The client sends topic, optional replacement and submission key. The route resolves the authenticated workspace and validates the bounded request.
- Resolve identity. The service loads configured, verified credentials, builds the normalized request hash and checks the scoped seven-day cache. A hit returns immediately.
- Approve the cold request. The server issues the US$0.25 signed challenge. The client resubmits it with the original key. The server verifies all bindings.
- Reserve before spending.
reserveProviderTask()atomically claims the attempt and quota slots. The service resolves that exact credential version and enterssubmitting. - Execute the dependency graph. Discover resources, normalize/deduplicate candidates, verify status, enrich up to five failed destinations, then parse up to ten referring domains.
- Derive the internal result. Normalize and deduplicate provider rows, associate contact evidence, calculate relevance/score and retain the top fifty opportunities.
- Persist and settle. Save rows, coverage counters, warnings and capture time in
provider_cache, settle task/quota state and record the usage receipt. - Join replacement evidence. The route reads the replacement site’s saved crawl, ranks eligible pages and adds suggestions plus any manual URL. This happens outside the paid discovery cache.
- Render and reuse. Return
DeadLinkOpportunitiesVM. React renders rows, warnings, contact provenance and replacement choices. Matching future requests reuse provider evidence; replacement changes use storage-only enrichment.
Further implementation
Metric provenance, crawl behavior, contracts and delivery details.
Deep implementation detailsMetrics, units and financial interpretation+
| Metric | Origin and calculation | Meaning |
|---|---|---|
| Search volume | Provider keyword metrics, mapped to monthlyVolume. | Estimated search demand. It is not visits to the user’s site. |
| CPC | Provider cpc, preserved as a nullable numeric metric. | Estimated advertising cost per click. Not customer value, profit or lifetime value. |
| Estimated traffic / ETV | Organic mapper retains standard etv and optional clickstream_etv, selecting max(standard, clickstream ?? 0) and recording the selected method. | Estimated traffic volume, not dollars or actual Google Search Console clicks. Missing standard ETV currently falls back to zero in this mapper. |
| Opportunity score | The explicit dead-link or keyword heuristic above. | Prioritization points, not revenue, ROI or a success probability. |
| Provider spend | Server estimate and returned provider cost, with nullable actual cost. | Cost of obtaining research evidence, not the economic value of the opportunity. |
The reviewed hosted flows do not calculate customer LTV, conversion-based opportunity revenue or a user-entered customer-value model. DataForSEO also exposes an estimated paid-traffic-cost metric, but the active organic bundle does not map that into its report. I would not relabel traffic volume as financial value. Provider metric definitions ↗
Implementationsrc/server/providers/dataforseo/hosted-organic-bundle.ts · metricsFromItem() src/server/services/research-service.ts · buildOpportunity() src/lib/provider-cost.ts
Deep implementation detailsCrawl lifecycle and Site Mode+
- StartPOST technical · action:start
- SubmitOnPage task_post → persist remote ID
- AdvancePOST technical · action:collect
- Pollsummary/{id} → ready?
- Collectpages + links + non_indexable
- CommitSnapshot → pages → issue occurrences
/api/explore/technical?target=… uses POST actions start and collect. GET reads the saved report without starting or polling provider work. Start accepts a page limit of 25, 100, 250 or 500 and options for subdomains, sitemap, JavaScript and resources. Defaults are 250 pages, same-domain scope, sitemap enabled, JavaScript/resources disabled.
The adapter posts target, max_crawl_pages, allow_subdomains, respect_sitemap, load_resources, enable_javascript and store_raw_html:false to /v3/on_page/task_post. The returned ID is persisted in provider_tasks.provider_task_reference, linked from the technical run. A later collect action polls /v3/on_page/summary/{id}, with a five-minute next-poll interval.
Once ready, a fifteen-minute collection lease coordinates advancement. collectHostedOnPageAudit() requests /v3/on_page/pages, /v3/on_page/links and /v3/on_page/non_indexable concurrently, each with {id, limit: pageLimit}. These are bounded collections, not exhaustive pagination of every provider record. Mapped pages, categorized issues and per-page occurrence evidence become a saved snapshot. The previous completed snapshot remains readable while a new run is underway.
An incomplete provider task retains its remote reference for another advance request; runs have a two-hour age limit. Collection calls have a 120-second timeout and an 8 MiB response cap. A collection failure settles the attempt as failed rather than saving a partially assembled new snapshot. This is continuation of an asynchronous task, not arbitrary per-stage recovery for every feature.
| Keyword Mode | Site Mode | |
|---|---|---|
| Input | Topic + optional replacement site/URL | Domain with an existing technical snapshot; optional manual replacement |
| Service | requestHostedDeadLinkKeyword() | getHostedSiteDeadLinks() |
| Evidence | Fresh provider pipeline or matching provider cache | Latest workspace-owned crawl issues and occurrence JSONB |
| Provider work | Can initiate four kinds of paid provider request | GET initiates none; a separate technical scan creates evidence |
| Derivation | Weighted topic/source score; publisher contacts | Group by destination/source pair; preserve each repair location and count repeat occurrences |
| Priority | Weighted formula, up to 50 rows | Fixed 90 for internal repairs, 65 for external failures; relevance 100 |
Site Mode reads up to 200 issues, keeps BROKEN_INTERNAL_LINK/BROKEN_EXTERNAL_LINK, and pages through occurrences in batches of 200, up to 1,000 per issue with a truncation warning. It reads evidence.links or a single occurrence and groups by destination/source pair, preserving the pages where a repair is needed. The common DeadLinkOpportunityRowVM lets both modes use one renderer, but site rows leave source rank/spam/contact unknown and use a placeholder follow flag. They are repair priorities, not enriched outreach rankings. Both modes then use the storage-only replacement service; without a snapshot, the interface asks the user to run a scan.
Backlink Research is another bounded read model
The standalone backlink adapter requests /v3/backlinks/summary/live for aggregate metrics, then /v3/backlinks/backlinks/live for up to 100 rank-ordered rows. Summary requests live links with subdomains/indirect links included and internal links excluded; row requests include both live and lost status. HostedBacklinksEvidencePayload separates aggregate metrics from BacklinkEvidenceRowVM[], including URLs, anchors, ranks and first/last-seen timestamps. Both calls must succeed before saving this bundle; that differs from the dead-link pipeline’s optional enrichment behavior.
Implementationsrc/server/services/hosted-technical-audit-service.ts src/server/providers/dataforseo/hosted-onpage-audit.ts src/server/providers/dataforseo/hosted-backlinks-summary.ts hosted-dead-link-service.ts · getHostedSiteDeadLinks()
Deep implementation detailsSchema excerpts and failure contracts+
type DeadLinkOpportunityRowVM = Readonly<{
deadUrl: string;
statusCode: number | null;
sourcePageUrl: string;
sourceRank: number | null;
spamScore: number | null;
relevanceScore: number;
opportunityScore: number;
replacementUrl: string | null;
contact: DeadLinkContactRouteVM;
}>;For example, provider rank ?? domain_from_rank becomes one bounded sourceRank, and backlinks_spam_score ?? spam_score becomes spamScore. The UI consumes that stable row shape alongside an EvidenceState, counters and a usage receipt. It does not branch on provider endpoint field names. Drizzle’s inferred record types describe persistence; view models in src/lib/contracts/page-vms.ts describe the application interface.
Runtime validation is layered. Hosted routes allowlist keys and bound JSON size, mutations check origin/content type, and input normalizers validate domains and numeric limits. Signed confirmation uses a strict Zod schema. Provider adapters validate response envelopes, bound streamed bytes and map fields explicitly. This is not an assertion that all provider data passes through one universal Zod schema.
The keyword adapter expects a successful envelope, exactly one task, task status 20000 and a result array. It enforces a 25-second timeout, a 1,500 KiB response limit and JSON depth/node/array bounds. The mapped research payload is sanitized again before storage, including finite numbers, bounded strings, public URLs and rejection of prototype-related keys. Dead-link mapping checks successful task status too, then validates cached rows against 404/410 status, safe source/destination URLs, domain-associated email evidence and integer scores/counters.
| Condition | Behavior in the reviewed implementation |
|---|---|
| Metric omitted | Most keyword/rank/spam fields remain null or unknown. Exceptions matter: a present trend month with missing volume becomes zero; organic standard ETV falls back to zero. |
| Provider timeout or oversized response | Bounded adapter returns unavailable/unexpected-response failure. No automatic retry loop in the hosted live keyword/dead-link adapters. Possible provider spend is not assumed to be zero. |
| Discovery or verification fails | Dead-link execution fails. Later stages cannot run without their prerequisite evidence. |
| One backlink/contact enrichment fails | Continue with other evidence and add a warning. Cached view state becomes partial when warnings exist, even with usable rows. The response retains warning messages, not a checkpoint for restarting each failed call. |
| No email found | Null email, unknown kind, low confidence after parsing; unknown confidence for failed/unrequested parsing. The publisher homepage is a manual-research fallback, not an asserted contact form. |
| Ambiguous page status | Only 404/410 with successful provider task status qualifies. Unknown, 403, 429 and 5xx are not dead-link evidence. Missing checks are disclosed rather than called healthy. |
| Duplicate URLs / domain variants | Candidate URL Set removes exact normalized duplicates; lowercase/www host grouping caps fan-out. Destination/source Map removes repeated backlink pairs. Query and scheme variants can still remain distinct. |
| Spam source | Invalid referring URLs are dropped. Valid URLs with high spam scores remain eligible but receive less quality weight. |
| Empty or incomplete completed result | Dead-link view state is partial when rows are empty or warnings exist. That is a presentation state, not a database partial-completion state. |
| Stale evidence or duplicate submission | Cache expiry/credential identity controls reuse; snapshot reads retain capture context. Idempotency replays the same attempt, and stale reservations are reconciled conservatively. |
Implementationsrc/lib/contracts/page-vms.ts src/server/security/hosted-api.ts src/server/providers/dataforseo/hosted-keyword-research.ts src/server/services/hosted-keyword-research-service.ts
Deep implementation detailsReport exports and WebGL implementation+
Server-generated Excel workbooks
buildHostedExplorerWorkbook() uses ExcelJS behind /api/explore/export. It calls workspace-scoped report getters and flattens saved view models into worksheets for selected report categories, including organic, backlink and technical evidence. It does not start fresh provider research. The dead-link prospect list is not one of this export route’s report categories.
ExcelJS supplies multiple worksheets, typed cells, column widths, frozen headers, filters and styles. Export metadata records the target, market, generation time and selected report context. URLs, status fields and captured evidence travel as columns; absent reports produce an unavailable entry instead of fabricated values. The server bounds export rows at 12,000 and the resulting workbook at 4 MiB.
safeSpreadsheetText() normalizes text and neutralizes formula-leading characters before writing untrusted strings. Numeric and boolean values remain native cell values; null becomes an empty cell. The workbook is serialized on the server, keeping the workbook engine out of the browser’s report interaction path.
The website’s procedural WebGL visual
SilkShader draws two triangles with six vertices. A fragment shader calculates the surface from time, resolution, pointer and color uniforms, with the palette read from CSS design tokens. React selects the shader mode; animation updates uniforms and calls drawArrays(), rather than rebuilding geometry every frame.
The loop stops for pause, reduced motion, offscreen visibility and a hidden document. Pointer motion interpolates toward its target. The graphite mode gates rendering around 30 FPS; the component also caps its drawing buffer through its existing resolution policy. WebGL failure has a static fallback, context loss stops drawing and cleanup releases its resources. These are properties of the Silktrace website implementation, not performance changes to this portfolio’s 3D room.
Dead-link CSV preserves the reviewed choice
deadLinkCsv(vm, selected) is a separate client-side export. It flattens the in-memory report and each row’s selected replacement into CSV with mode, query, capture time, failed destination/status, referring page, anchor, occurrence count, priority, replacement provenance/overlap/matched terms, contact source and warnings. Site Mode omits the outreach priority value.
Values pass through safeSpreadsheetText(), quotes are escaped, and the file uses a UTF-8 byte-order mark and CRLF rows. A browser Blob/object URL downloads the file without another API request. This preserves what the operator actually selected without rerunning discovery or replacement matching.
Implementationsrc/server/services/hosted-explorer-export-service.ts src/lib/dead-link-export.ts src/components/marketing/shader.jsx
Scope a related build through custom web development, or explore how the studio supports B2B and SaaS teams.