From Amazon Seller Data to AI Answers: Building an MCP Data Pipeline for a Multi-Brand Platform

The Amazon Seller MCP Data Pipeline gives a multi-brand platform one authenticated way to answer seller questions from an AI client such as Claude. Sellers connect their Amazon account once, and the platform collects, archives, validates, and stores their orders, inventory, financials, and catalog data in an isolated per-brand database. Claude then reads that data through 22 read-only Model Context Protocol (MCP) tools, so “what sold last week?” no longer starts with a manual report download. Our engineering team built and verified the full path in about four weeks.

Our Purpose

  • Replace manual Amazon report exports and one-off scripts with a single question interface over each seller’s own data
  • Unify reports, API operations, and push notifications that arrive as TSV, XML, JSON, and CSV on different schedules
  • Keep every brand’s data separated at the database level, so one missing filter can never expose another brand
  • Give AI clients records from a known source and a known time period, every time they answer

How We Delivered

  • Built a NestJS 11 and TypeScript backend with a Python 3.11 worker, using Redis and BullMQ queues to schedule collection from every 15 minutes to daily
  • Designed AI-powered data pipelines that archive raw payloads to Amazon S3 first, then validate and upsert records so reruns never duplicate data
  • Gave each brand its own PostgreSQL schema and read-only role, following the same isolation thinking we use to build multi-tenant SaaS on AWS
  • Shipped a declarative migration tool with dry-run mode and an audit table, plus Docker images and Jenkins pipelines for deployment
Learn more

Game-Changing Features

  • 22 read-only MCP tools covering orders, inventory, financials, settlements, pricing, invoices, returns, performance, and sales analytics
  • Side-by-side comparisons across brands and marketplaces, with every query scoped to the caller’s own brands
  • Amazon SP-API OAuth flow with refresh tokens encrypted at rest using AES-256-GCM and short-lived authorization state
  • A source registry of 43 report types, 27 API operations, and 8 notification types that configuration can switch on in stages
Learn more

Values Achieved

  • Proved the full path from signup to AI-readable seller data in one recorded end-to-end session
  • Turned job retries from a hazard into a non-event, because every write is an idempotent upsert
  • Made every answer traceable, since the original Amazon payload in S3 can be replayed when a transformation rule changes
  • Gave Claude a question-only interface that can never change seller data or see another seller’s records

From seller authorization to scheduled collection, validated storage, and AI answers in Claude, all in one pipeline

Brand operators were answering every business question with a manual export or a one-off script, and no AI assistant could answer reliably without known, validated sources. The goal was a modular platform where a user signs up, connects their Amazon seller account once, and then simply asks their AI client. We combined generative AI integration over MCP with a cloud-native application backend on AWS, so the platform collects, validates, and stores the data behind the scenes for every brand.

Jumpstart My Project

Overcoming the Challenge — Fragmented Formats, Overlapping Data, Shared Risk

Amazon delivers seller data through document reports, API operations, and push notifications, each with its own format, schedule, marketplace, and permission rules. The same order can arrive from two sources, and a report can be delivered twice, so careless ingestion produces duplicates and conflicting numbers. Many brands share one platform, where a shared-table design sits one missing filter away from a cross-brand leak. Amazon calls are also rate limited and fail from time to time, so collection had to be retryable and resumable from the first day.

👏🏽Transformative Solution — Archive First, Isolate Every Brand, Read Only

The worker stores every raw payload and report document in S3 before transformation, so a wrong rule means a replay instead of a re-fetch from Amazon. A separate job normalizes money to decimals, timestamps to UTC, and country and currency codes, then upserts into the brand's own PostgreSQL schema. Progress markers and backoff retries let each run continue where the last one ended. Claude reaches that data only through read-only MCP tools checked against a hashed bearer key, so it can ask questions but never change seller data.

The Outcome

A recorded end-to-end run covered signup, login, brand creation, seller authorization, collection, transformation, MCP reads, and the readiness email in a single session, with database queries confirming the stored records afterwards. Sellers now connect Claude with one endpoint and one key, and Claude picks the right read-only tool for each question on its own. Every answer comes from validated records in the brand’s own schema, and the original payload in S3 lets anyone trace a number back to its source.

Lets Talk

ONE AI QUESTION INTERFACE

REPLAYABLE RAW DATA

SAFE, REPEATABLE WRITES

PER-BRAND DATA ISOLATION

Tech Stack

Some technologies used for this project