Similarity API
Compares records field-by-field (numeric closeness and string edit distance) and returns pairs above a similarity threshold. Use it to surface fuzzy duplicates before merging.
What it does
Score pairwise similarity within one record set.
When to use it
Reach for Similarity when you need to answer: “Which records in this set look like near-duplicates?” It sits in the Reconcile stage of the catalog — match expected against actual.
Endpoint & cost
POST /api/v1/similarity
Cost: Per credit. Zaps draw down your shared monthly allowance; requests that fail validation are not charged.
Authentication
Send your key as a bearer token: Authorization: Bearer zap_live_…. See Authentication for key management and the alternate header.
Example request
curl https://zapinner.com/api/v1/similarity \
-H "Authorization: Bearer zap_live_••••••" \
-H "Content-Type: application/json" \
-d '{"data":[{"name":"Acme Inc"},{"name":"ACME, Inc."},{"name":"Globex"}],"options":{"threshold":0.6}}'request_id. Full field-level request and response schemas for this endpoint live in the OpenAPI spec and the API reference.Errors & limits
Errors use one consistent shape ( error.code, error.message, request_id). See Errors and Rate limits.
SDK
Similarity has typed methods in both the JavaScript and Python SDKs. Both are implemented and package-ready but not yet published to npm / PyPI — until they ship, vendor the package source or install from a git ref, or call the endpoint over HTTPS with the example above.
Try it
Select Similarity in the playground to see the request, or run it live with your API key.
Related APIs
- Reconciliation — Compare two datasets and surface mismatches, duplicates, and missing records.
- Entity Resolution — Cluster records into entities and emit one canonical record each.
- Conflicts — Find fields whose values disagree across records sharing a key.
- Merge — Collapse duplicate records into one golden record with provenance.
- Survivorship — Collapse keyed groups into golden records with survivorship rules.
- Link — Record linkage across two datasets.
