[No. 001]Web Scraping

AI Tools Hub

Sophisticated AI Tool Directory

SR
bySanthosh Reddy
RoleFull Stack Engineer
TimelineNovemver 2025 - December 2025
Read4 min
AI Tools Hub
FIG. 01 - AI Tools Hub overviewAI Tools Hub.case

Introduction

AI Tool Hub started as a question: what would a discovery platform look like if it autonomously crawled, deduplicated, and categorized the entire ecosystem of new AI tools in real time? The product is a centralized intelligence hub featuring an automated multi-source scraping orchestrator, LLM-powered categorization, and intelligent cross-domain deduplication, deployed natively across web and Android.

The problem with manual directories

Existing AI tool directories rely heavily on manual submissions or static, curated lists. They cannot autonomously ingest tools from disparate global sources, do not inherently resolve duplicated entries across different aggregators, and lack a unified method for organizing rapidly expanding categories. For someone trying to stay on the bleeding edge of AI, that is the inverse of useful.

A discovery platform should crawl the internet, deduplicate the noise, and structure the intelligence autonomously before a user ever searches for it. Anything less is just an outdated bookmark manager.

Built with
React & TypeScriptViteSupabase (PostgreSQL, Edge Functions)Tailwind CSSCapacitorLLM APIs

Architecture

I engineered a fully autonomous, hourly data collection pipeline managed by a centralized scrape-orchestrator Supabase Edge Function. This orchestrator delegates tasks to specialized worker functions that extract real-time updates from disparate sources including GitHub, HackerNews, HuggingFace, ProductHunt, and YouTube. On top of this sits an intelligent ingestion layer leveraging fuzzy name matching and domain comparison to mathematically guarantee zero duplication across the ecosystem.

The data normalization process is powered by a dedicated LLM utility that autonomously categorizes incoming tools and standardizes their metadata before insertion. The resulting structured intelligence is indexed in PostgreSQL using custom full-text search RPC functions, exposing sub-millisecond query capabilities to the React frontend and the Capacitor-compiled Android application simultaneously.

What I prioritized

A few of the technical decisions that mattered most:

  • Autonomous Ingestion. A fleet of Supabase Edge Functions orchestrated to scrape multiple distinct platforms simultaneously on a strict hourly cron schedule.
  • Intelligent Deduplication. Algorithmic URL matching, fuzzy logic, and domain cross-referencing to eliminate redundant tools fetched from overlapping aggregators.
  • LLM-Powered Categorization. Leveraging AI to automatically read tool descriptions and sort them into a standard taxonomy without human intervention.
  • Cross-Platform Native Deployment. Utilizing Capacitor to wrap the responsive React/Tailwind web interface into a native Android APK from a single codebase.

Where it goes next

The current roadmap is focused on refining the scraping logic, expanding the sources, and enhancing the search algorithms with vector embeddings for semantic discovery. The longer-term aim is moving beyond a directory to become an active recommendation engine that pairs users with the exact AI tool they need based on their specific workflow constraints.