This system integrates an open-source web scraping engine deployed via Docker with a no-code workflow that automates website crawling, content extraction, and data ingestion for retrieval-augmented generation. The workflow uses HTTP requests to trigger asynchronous scraping tasks, parses sitemap.xml for URLs, processes markdown and HTML responses, and feeds vector embeddings (via OpenAI) into a Supabase vector store. The process also includes AI agent components such as Pydantic AI and TEN Agent.
You'll bring your own API credentials (so the workflow runs under your accounts) — I'll walk you through wiring them up on our setup call, which is included with your purchase.
Schedule a free call and I'll help you identify the highest-impact automation opportunities for your business.
Book a Call