Dedupe API

Detect records that represent the same entity within a single dataset, with a transparent similarity basis.

What it does

Find duplicate records within a dataset.

When to use it

Reach for Dedupe when you need to answer: Which records are duplicates? It sits in the Prepare stage of the catalog — clean and de-duplicate raw data.

Endpoint & cost

POST /api/v1/dedupe

Cost: 1 credit. Zaps draw down your shared monthly allowance; requests that fail validation are not charged.

Authentication

Send your key as a bearer token: Authorization: Bearer zap_live_…. See Authentication for key management and the alternate header.

Example request

curl https://zapinner.com/api/v1/dedupe \
  -H "Authorization: Bearer zap_live_••••••" \
  -H "Content-Type: application/json" \
  -d '{"records":[{"id":1,"email":"a@acme.com"},{"id":2,"email":"A@ACME.com"}]}'
Every response is structured JSON and includes a request_id. Full field-level request and response schemas for this endpoint live in the OpenAPI spec and the API reference.

Errors & limits

Errors use one consistent shape ( error.code, error.message, request_id). See Errors and Rate limits.

SDK

Dedupe has typed methods in both the JavaScript and Python SDKs. Both are implemented and package-ready but not yet published to npm / PyPI — until they ship, vendor the package source or install from a git ref, or call the endpoint over HTTPS with the example above.

Try it

Select Dedupe in the playground to see the request, or run it live with your API key.

  • Repair Send messy data. Get production-ready data back.
  • Map Turn incoming data into the exact schema your application needs.
  • Guard Validate, repair, or reject data before it reaches your application.
  • Normalize Turn messy, inconsistent fields into clean, structured, usable data.
  • Transform Apply declarative field operations to reshape records.
  • Validate Check records against a declared field ruleset.
  • Standardize Coerce every value into the canonical form for its column's type.
  • Parse Parse declared string fields into typed values and name components.
  • Type Inference Infer each column's data type by majority vote over its values.
  • Schema Map Map source records onto a target schema with per-field confidence.
  • Redact Remove detected PII from string fields, replaced with placeholders.
  • Mask Mask detected PII while preserving recognizable shape.
  • Impute Fill missing values with a chosen strategy.
  • Normalize Phone Normalize phone numbers to E.164 format.
  • Normalize Address Canonicalize address strings to a consistent form.
  • Normalize URL Canonicalize URLs (scheme, host, port, path, query).
  • Text Clean Normalize whitespace, case, and punctuation in text.
  • Cluster Group records into clusters by similarity.
  • Rank Weighted-score and rank records by chosen fields.
  • Segment Assign records to segments with first-match rules.
  • JSON Repair Repair common JSON syntax errors into valid JSON.
  • Flatten Flatten nested records into dot-path keys.
  • Unflatten Rebuild nested records from dot-path keys.
  • Convert Convert records between JSON and CSV/TSV.