goldenmatch
Eliminates manual deduplication and record matching work that typically consumes hours per data cleanup cycle.
Entity resolution library for Python and TypeScript that dedupes and matches messy records into golden records with zero configuration — fuzzy, exact, probabilistic (Fellegi-Sunter), and LLM matching. Scores 96.4% F1 on the DBLP-ACM benchmark and scales from laptop CSVs to 100M+ rows on a Ray cluster. Includes an MCP server plus a wider data-quality suite (GoldenCheck, GoldenFlow, GoldenPipe, InferMap).
- Merge duplicate customer records before loading into CDP
- Deduplicate vendor master data across procurement systems
- Resolve patient records across hospital EHR systems
Eliminates manual deduplication and record matching work that typically consumes hours per data cleanup cycle. Reduces downstream data quality issues that corrupt reporting and customer analytics.
Data engineering and analytics teams managing CRM, customer database, or transactional records with duplicates and inconsistent entity references across systems.
https://github.com/benseverndev-oss/goldenmatch
By benzsevern
How to Get It
pip install goldenmatch
Tip: Paste this into a Claude Code conversation. Verify command matches your Claude Code version.
After installing, paste this into Claude:
Help me merge duplicate customer records before loading into CDP
Trust Signals Auto-scanned
Community Pulse Emerging
Discussed on Hacker News, Reddit
- Show HN: GoldenMatch – Entity resolution with LLM scoring, 97% F1, no Spark — Hacker News · 3 pts
- Show HN: GoldenMatch – 100M-row dedupe on Ray in 213s, no Spark, Arrow-native — Hacker News · 3 pts
- GoldenMatch – Find duplicate records in 30 seconds. Zero-config entity resolutio — Reddit · 1 pts
3 mentions across 2 sources
Reviewer notes
Auto-scanned review. These are observations, not a security certification.
catalog_hygiene stale-eval refresh: Scored from trust signals (evidence-eval-v1): 128 GitHub stars; 4 contributors; last commit 0d ago; license MIT.
Things to check
- Scanned, not hands-on tested — this entry was auto-scanned from public metadata (GitHub metrics, license, security flags). No reviewer has run it, and no tool-specific limitations have been documented yet.
How to evaluate tools before deploying →
Data shown here comes from public APIs and automated scanning. Reviewer notes reflect one person's experience. This is not a security certification or legal recommendation. Always evaluate tools according to your own organization's policies.
Evaluation
catalog_hygiene stale-eval refresh: Scored from trust signals (evidence-eval-v1): 128 GitHub stars; 4 contributors; last commit 0d ago; license MIT.