Ingest, classify, enrich, and publish hotels, attractions, and transfers across any destination — powered by Claude AI, Ahrefs keyword data, and a governed human review pipeline.
Pipeline Stages
Product Categories
Human Review Gates
Audit Trail
Every product follows a governed, auditable journey from raw data source to live catalogue entry.
Ahrefs keyword research + source discovery + approved website list compilation before ingestion begins.
Official APIs (Booking.com, Viator, GetYourGuide) as primary sources. Governed scraping only where ToS explicitly permits.
Every record stored with source URL, timestamp, confidence score, and ToS approval flag. Full raw audit trail.
Claude API classifies each record into Hotels, Attractions, Transfers, or Restaurants with a 0.0–1.0 confidence score.
Category-specific schema enforcement. Incomplete records flagged and routed to enrichment queue automatically.
Claude API + Ahrefs keywords generate descriptions, meta tags, FAQs, Schema.org markup, and image alt text per product.
Queue A: classification review. Queue B: content approval. Nothing reaches staging without passing both gates.
Full production mirror. Every batch validated before the Product Manager approves the production push.
Upsert logic with full rollback by job ID. Package Builder assembles packages automatically post-push.
Before fetching a single record, the system queries Ahrefs for the top 50 keywords for your destination, scores every candidate data source by domain authority and keyword relevance, and compiles a ranked Approved Source List — automatically, every run.
50 keywords researched per destinationOfficial APIs are the mandatory primary source. Web scraping is permitted only for sources that pass a legal approval checklist — with ToS status, robots.txt check, and data usage rights verified for every candidate source before it enters the pipeline.
0 unauthorised scraping targetsClaude API generates 7 content outputs per product: short description, long description, meta title, meta description, FAQ section (3–5 Q&As targeting Google PAA), Schema.org JSON-LD markup, and SEO image alt text — all keyword-grounded via Ahrefs.
7 SEO outputs generated per product7 package types with configurable JSON combination rules. Automated pricing engine: base cost + configurable margin + floor price guard. AI-generated day-by-day itineraries. Packages flagged as incomplete if any net rate data is missing.
20% default margin, configurable per type5 dimensions · 40+ tags · AI-suggested · Human-confirmed
11 modules. 5 user roles. Zero context-switching.
Trigger runs, monitor pipeline status
Real-time job progress per source
Classify low-confidence records manually
Approve or edit AI-generated content
Browse all products by destination/category
Edit fields, flag missing data, bulk edit
Assign Source 1/2/3 per product
Apply taxonomy tags, review AI suggestions
Assemble packages with pricing engine
Review diff, approve production push
Freshness heatmap, API health, errors
API-first policy. Scraping limited to verified ToS-compliant sources only.
Mandatory Human Review Queue A before any API push.
Full package typology + pricing engine defined.
Two mandatory queues: Classification + Content.
Managed scraping via Apify/Bright Data — not in-house.
Sign in to the platform or contact the Strategy & Analytics team to request access.