Full Stack2026
Dealify
Multi-Site Product Price Comparison & Availability Tracker
Overview
A full-stack product price comparison engine that aggregates data from Amazon, Flipkart, Croma, Reliance Digital and more. Provides live pricing, stock status, and historical price trends.
Architecture
Distributed scraper workers (Puppeteer, Playwright) feed into a normalization pipeline. REST + GraphQL API serves the React frontend. Redis with TTL invalidation handles caching; cron jobs refresh stale data. AWS EC2 + S3 backend, Netlify frontend.
Tech Stack
Node.jsTypeScriptGraphQLPuppeteerPlaywrightRedisDockerAWSReactNetlify
Challenges
- ▸Reverse-engineering undocumented site structures for resilient product lookups
- ▸Unified data normalization across heterogeneous e-commerce schemas
- ▸Rotating user-agents and rate-limiting to avoid scraper detection
- ▸Redis TTL strategy for freshness vs. latency trade-off
Lessons Learned
GraphQL is excellent for flexible aggregation queries across multiple data sources
Docker-encapsulated scraper workers make scaling and isolation easy
Cache invalidation strategy is as important as caching itself
Background cron jobs for data freshness decouple scraping from request latency