shipwithjev

Blog / Recipes / FIG. 75

Dedupe Your CRM in a Weekend

A weekend CRM deduplication recipe: block candidates, ask Jev if two contacts are the same person, review the unsure pairs, then merge safely.

Your CRM has three records for the same person: "Jon Smith, Acme", "jonathan.smith@acme.io", and "J. Smith (Acme Corp) DO NOT CONTACT". Exact-match rules missed all of them. This recipe handles CRM deduplication with Jev, TypeSafe AI's decision model, as a Saturday-to-Sunday project with a human on the merge button.

The theory (pairwise verdicts, why blocking matters, identity linking in general) lives in entity resolution. This page is the weekend plan.

Saturday morning: export and block

  1. Export contacts to a working table. Name, email, company, phone, title, and the record ID. Work on a copy. Never run dedup experiments against the live CRM.

  2. Normalize the boring stuff first. Lowercase emails, strip whitespace, standardize phone formats, remove "Inc." and "Ltd." from company names. Deterministic cleanup catches the easy duplicates for free, and you don't need a model to tell you two identical emails match.

  3. Generate candidate pairs with blocking. Comparing every contact to every other contact is the quadratic trap that entity resolution explains. Instead, only pair records that share something cheap: same email domain, same last name initial plus company, same phone suffix. A 20,000-row CRM should produce thousands of candidate pairs, not hundreds of millions.

Saturday afternoon: fuzzy match with verdicts

  1. Write the pairwise question. Something like: "Do these two CRM records refer to the same real person? YES / NO / UNCLEAR." Pass both records' fields side by side. Include the escape hatch; "UNCLEAR" is where the interesting cases hide.

  2. Run the pairs through Jev. Per ecosystem documentation, the model is typesafe-ai/jev via the Vercel AI Gateway, returning a probability per choice. The official syntax is at docs.typesafe.ai.

# pseudocode, not real API syntax
for (a, b) in candidate_pairs:
    v = jev.choose(SAME_PERSON_Q, format_pair(a, b), ["YES", "NO", "UNCLEAR"])
    save(a.id, b.id, v.top_choice, v.top_probability)

If your contacts already live in Postgres, there's a shortcut worth knowing: a community extension runs Jev as a filter inside SQL (build), and natural-language database queries covers how it works and its reported throughput. Pairwise matching still needs the pair table, but the verdict column can live in the database.

Sunday: review, then merge

  1. Bucket by confidence. Three piles:

    • High-confidence YES: probable duplicates, queued for merge.
    • High-confidence NO: done.
    • Everything else: human review.
  2. Spot-check the YES pile. Open 50 random high-confidence matches and check them. If more than a couple are wrong, tighten the question (for example, require matching company or matching email local part) and rerun. Reruns are cheap; bad merges are not.

  3. Review the unsure pile by hand. This is usually the smallest pile and the most important. Father and son at the same company, two people named Priya Patel at one enterprise, a personal and work email for the same buyer. A person resolves these in seconds; a model shouldn't be guessing.

  4. Merge in batches, with an undo plan. Merging records is effectively irreversible in most CRMs: activity history, owners, and deal links get combined. So a Jev YES is never the merge trigger by itself. A human approves each batch, you export a pre-merge snapshot, and you merge in chunks of a few hundred so a mistake stays small.

Why not just use the CRM's built-in dedup?

Fair question, and you should run it first. Built-in tools are great at exact and near-exact matches. They struggle with the semantic ones: nicknames, rebrands, "same person, new job", initials. That residue is what this weekend is for.

The other honest objection: "Couldn't a chat model do this?" It could, at chat-model prices and with free-text answers you'd have to parse. A closed YES/NO/UNCLEAR with a probability is the natural format for match decisions, which is why decision models fit here.

Keeping it clean after the weekend

Add the same question as a check on new record creation, so duplicates get flagged before they multiply. The webhook pattern in Jev automations does this without a custom service. Flag, don't auto-merge.

Frequently asked questions

How long does it take to dedupe a CRM with Jev?

The verdict pass itself is fast; the time goes into export, blocking, and human review, which is why this fits a weekend. Plan most of Sunday for reviewing the unsure pairs.

Should Jev merge duplicate contacts automatically?

No. Merges are hard to undo, so a human should approve every batch, with a snapshot taken first. Jev's job is ranking which pairs are worth a look.

Do I need blocking for a small CRM?

For a few hundred contacts you can compare everything, but past a few thousand the pair count explodes. Entity resolution explains the math and the blocking strategies.

Can I run the matching inside my database?

If your data is in Postgres, a community extension brings Jev verdicts into SQL. See natural-language database queries for the pattern and its limits.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.