● Context
Most TAM exercises end with a spreadsheet of 50,000 accounts dumped into the CRM. Six months later, nobody has called any of them. The file sits in a shared drive, last opened in January, while sales reps keep prospecting the same 200 accounts they found on LinkedIn.
We built Vizzia’s entire total addressable market from scratch last year: 35,000 municipality records covering 98.3% of French communes, enriched with SIRET numbers, emails, and phone contacts. It took four weeks, not four months, and the database went live inside HubSpot on day one. That project taught us more about TAM building for B2B than any framework slide deck ever could.
This guide walks through the exact process we use at Cashmyrr to build operational TAMs that sales teams actually work from. Not vanity market sizing for pitch decks. Actual prospect databases that generate pipeline.
What is TAM, SAM, SOM?
These three acronyms get thrown around in board meetings and fundraising decks, but the definitions matter for operational purposes too.
TAM (Total Addressable Market) is the entire universe of companies or organizations you could theoretically sell to if you had unlimited resources, distribution, and product coverage. For a CRM consulting firm, that might be every B2B company with more than 10 employees in Europe. For a public-sector SaaS product, it might be every municipality in France.
SAM (Serviceable Addressable Market) is the slice of the TAM you can actually reach with your current product, pricing, and go-to-market motion. If your product only supports French-language onboarding and your sales team is based in Paris, your SAM is a lot smaller than your TAM regardless of what your product could do in theory.
SOM (Serviceable Obtainable Market) is what you can realistically close in the next 12 months given your current team size, pipeline capacity, and sales velocity. If your average deal cycle is 90 days and you have three AEs, your SOM has a ceiling no matter how large your SAM is.
The relationship between these three layers is what makes TAM exercises useful. A massive TAM with a tiny SOM tells you something about your go-to-market constraints. A SAM that closely matches your TAM means you’re either very focused or underestimating the filtering you should apply.
For B2B companies, the bottom-up TAM (counting actual companies) matters far more than the top-down version (starting with total market revenue and dividing). We’ll get to that distinction later.
Why Most TAM Builds Fail
We’ve inherited enough broken TAM projects to spot the patterns. Here’s where they go wrong.
The “dump 50K accounts” approach. Someone buys a ZoomInfo or Apollo export, filters by industry and headcount, and uploads everything into the CRM. Sales gets a territory with 3,000 accounts, no context on which ones matter. Reps cherry-pick 50 familiar names and ignore the rest.
No ICP alignment. The TAM was built without involving sales leadership. Marketing defined the criteria based on ad targeting parameters. Sales looked at the list and said, “Half of these are too small to buy from us.” Firmographic filters alone don’t capture buying signals or organizational fit.
Single data source. Every database has gaps. Apollo has strong US tech coverage but weaker European SMB data. LinkedIn gives you people but not company financials. Government registries have legal entity data but no contacts. A TAM built from one source inherits all of that source’s blind spots.
No deduplication against existing CRM records. You import 20,000 accounts and create 4,000 duplicates because nobody checked what was already in the system. Now your sales team has two records for the same company with different owners and conflicting notes. Your CRM audit will catch these eventually, but the damage to rep trust is immediate.
No refresh policy. B2B data decays at roughly 22.5% per year according to Gartner. People change jobs, companies merge, phone numbers go dead, email domains change. A TAM built in Q1 is already 5-6% stale by Q2. Without a refresh schedule, your “comprehensive” database becomes a graveyard of bounced emails and disconnected numbers.
No segmentation or tiering. Treating every account in your TAM equally is a resource allocation failure. A 500-person company with a known pain point is not the same as a 15-person company in a tangentially related industry. Without tiers, you’re asking sales to prioritize randomly.
● Explanations
The 7-Step TAM Build Process
Here’s the process we follow for every TAM engagement. The order matters because each step depends on the outputs of the one before it.
Step 1: Define Your ICP with Stakeholders
This is a cross-functional exercise. Get your CEO, head of sales, head of CS, and marketing lead in the same room for 90 minutes. If you have a RevOps function, they should facilitate.
Work through these dimensions:
- Industry/vertical. Be specific. “SaaS” is too broad. “B2B SaaS selling to mid-market in HR tech” is useful.
- Company size. Revenue is better than headcount when you can get it. A 50-person consulting firm and a 50-person VC-backed startup have very different budgets.
- Geography. Where can you actually sell and deliver? Be honest about language, timezone, and regulatory constraints.
- Tech stack signals. Do your best customers use HubSpot? Salesforce? SAP? Tool usage can be a strong proxy for budget and sophistication.
- Budget indicators. Recent funding rounds, job postings for relevant roles, existing vendor relationships visible on LinkedIn or G2 reviews.
- Negative filters. Who do you explicitly not want? Government agencies, companies below a revenue threshold, specific sub-industries that look similar but have different buying processes.
Write this up as a one-page ICP document. Get sign-off from sales leadership before proceeding. Skip this step and everything downstream is guesswork.
Step 2: Identify Authoritative Data Sources
Once you know who you’re looking for, figure out where that data lives. This varies dramatically by market and vertical.
Government registries are the most reliable source for company existence and legal data. In France, INSEE provides SIRET numbers, activity codes (NAF/APE), and registered addresses. In the UK, Companies House offers similar data through a free API. These sources lack contact information but confirm that a company exists and where it operates.
Industry associations and trade directories often maintain member lists that serve as excellent vertical-specific TAM sources. The French Association des Maires de France publishes elected officials’ data. Professional federations in construction, healthcare, or consulting maintain directories with contact information.
Commercial databases like Apollo, ZoomInfo, Cognism, and Lusha fill the enrichment layer: emails, phone numbers, tech stack data, org charts. Each has different strengths depending on geography and company size. We compared the major enrichment tools in a separate piece.
Web scraping targets cover the gaps. Company websites, job boards (for hiring signals), review platforms like G2 or Capterra (for tech stack intelligence), and industry-specific platforms all contain structured data if you know how to extract it.
For any given TAM build, plan on combining at least two to three sources. One authoritative registry for company data, one commercial database for contacts, and one scraped source for specific signals your ICP requires.
Step 3: Validate with a Sample Scrape
Before you commit to a full build, run a sample of 500 to 1,000 records through the entire pipeline. This catches problems early.
Check these metrics on your sample:
- Match rate. What percentage of records from your primary source can be found in your enrichment tool?
- Field completeness. What percentage of records have a valid email? A phone number? A company size?
- Accuracy. Manually verify 50 records. Are the emails current? Are the companies still in business? Are the industry codes correct?
- ICP fit. Of your 500 records, how many actually match your ICP after enrichment? If it’s less than 60%, your source selection or filtering needs work.
On the Vizzia project, our sample scrape of 800 commune records revealed that the primary government directory had phone numbers for only 12% of entries. That finding changed our enrichment strategy before we scaled: we added a second source specifically for municipal contact data, which brought phone coverage to 31% at scale.
Step 4: Execute Full Scrape and Enrichment
With a validated sample, you scale. The tooling choice depends on the complexity of your sources and the volume of records.
For scraping: Apify handles JavaScript-heavy websites and has pre-built scrapers for common platforms. PhantomBuster is strong for LinkedIn-adjacent data extraction. For custom government registries, Python scripts with BeautifulSoup or Playwright handle authentication and pagination.
For enrichment: Apollo gives you up to 10,000 free credits per month and strong email finding. FullEnrich aggregates 15+ waterfall sources for phone numbers and emails. Cognism is the strongest option for GDPR-compliant European mobile numbers. Combine them to maximize coverage.
For orchestration: n8n ties everything together. A typical pipeline: trigger on new batch of records, call Apollo for email enrichment, fall back to FullEnrich for misses, validate deliverability via MillionVerifier, write results to a staging sheet, flag records needing manual review.
Run enrichment in batches of 500 to 1,000 to avoid API rate limits and monitor quality as you go.
Step 5: Clean and Deduplicate
Raw data is never CRM-ready. Budget real time for this step.
Standardize company names. “IBM” and “International Business Machines Corp.” are the same company. Use legal entity identifiers (SIRET, registration numbers) for exact matching and fuzzy string matching for the rest.
Normalize fields. Phone numbers in E.164 format. Countries in ISO 3166 codes. Industry classifications mapped to your CRM’s picklist values, not whatever the source provided.
Deduplicate against your existing CRM. Export current accounts and contacts, match incoming records by domain, registration number, or name-plus-city combinations. For fuzzy matching at scale, Dedupely works inside HubSpot. For more complex matching, LLM-powered deduplication through Claude or GPT-4 can evaluate whether “Mairie de Lyon” and “Ville de Lyon - Services Techniques” should be merged.
Flag, don’t delete. Mark potential duplicates for human review instead of auto-merging. Automated merges that go wrong create bigger problems than the duplicates themselves.
Step 6: Import into CRM with Segmentation
Do not import a flat list. Segment and tier before anything touches your CRM.
Tier accounts based on revenue potential and fit:
- Tier A: Perfect ICP fit, strong buying signals, high potential deal size. These get dedicated outbound sequences and named account owners.
- Tier B: Good ICP fit, moderate signals. These go into automated nurture sequences with periodic manual review.
- Tier C: Marginal fit or insufficient data. These stay in the database for marketing campaigns but don’t get direct sales attention.
Assign ownership based on territory, vertical, or round-robin depending on your sales model. An unowned account is an unworked account.
Set up CRM views and filters so reps can actually navigate their territory. A view showing “My Tier A accounts with no activity in 30 days” is worth more than a dashboard with aggregate TAM statistics.
Create a dedicated import record in your CRM with the source, date, and count. When someone asks “where did these 8,000 accounts come from?” six months from now, you want a clear answer.
Step 7: Set Up Refresh and Feedback Loops
A TAM is a living asset. Build the maintenance in from day one.
Quarterly data refresh. Re-run enrichment every 90 days. Flag changes: new emails, updated phones, companies acquired or shut down. At 22.5% annual decay, skipping one quarter means hundreds of stale records accumulating silently.
Sales feedback loops. Create a simple mechanism for reps to flag bad records: wrong industry, company too small, contact gone. This feedback should flow back to your data team and improve future enrichment runs.
Remove bad matches. If an account has been contacted three times with no response and research confirms bad fit, pull it from active territories. A smaller, cleaner TAM outperforms a bloated one.
Track coverage metrics. Monitor the percentage of your TAM with valid emails, phone numbers, and assigned owners. If email coverage drops below 70%, trigger a re-enrichment cycle.
● Conclusion
Top-Down vs Bottom-Up Calculation
Both approaches have their place, but they solve different problems.
Top-down starts with total market size data and narrows down. You take an analyst report that says the European HR tech market is worth $4.2B, estimate that your segment is 15% of that ($630M), and calculate how many customers at your average deal size make up that revenue. Useful for investor decks. It tells you the ceiling, not which companies to call on Monday.
Bottom-up starts with actual company records. You count every organization that matches your ICP, estimate revenue potential per account, and sum it up. The result is always smaller than the top-down number, but it’s grounded in real entities you can prospect.
For operational TAM building, always use bottom-up. A bottom-up TAM of 8,000 verified accounts with contact data beats a top-down slide claiming your TAM is $2 billion. The top-down number gives your board confidence in market size. The bottom-up database gives your sales team something to work with.
Use top-down as a sanity check. If your bottom-up TAM suggests 8,000 accounts worth $400M in aggregate, and an analyst report sizes the total market at $500M, your coverage looks reasonable. If the report says $50B, you’re either in a tiny niche or your ICP is too narrow.
Real Example: Building a 35K-Record Municipal Database
When Vizzia came to us, their target market was French municipal governments. No commercial B2B database covers elected officials and municipal decision-makers in France. Apollo and ZoomInfo barely index public-sector contacts. LinkedIn coverage for small-town mayors is thin.
Their existing CRM had 6,000 records with 6% SIRET coverage and roughly 1,000 phone numbers across the entire database.
Here’s what the build looked like:
Sources: We scraped the French government’s official directory of communes and municipal officials, cross-referenced with INSEE data for SIRET numbers and population statistics, and enriched with contact data from departmental association directories.
Scale: 35,000 commune records, covering 98.3% of all French communes. The missing 1.7% were mostly overseas collectivities with non-standard administrative structures.
Enrichment results: SIRET coverage went from 6% to 95.2%. Phone numbers went from 1,000 to 2,500. Email coverage reached 78% across the full database.
CRM integration: Everything was imported into HubSpot with tiering based on commune population, geographic territory assignment for Vizzia’s sales reps, and automated views filtering by department and region.
The entire project ran four weeks from kickoff to CRM go-live. Vizzia’s sales team started prospecting from the new database in week five. You can read the full case study here.
What made this work was the process, not the technology: precise ICP, right sources (government data over commercial databases), sample validation, aggressive cleaning, and segmentation applied before import.
● FAQ
For a straightforward B2B TAM using commercial data sources (Apollo, ZoomInfo) with light enrichment, expect 2 to 3 weeks including ICP definition, data extraction, cleaning, and CRM import. For complex builds requiring custom scraping, multiple source reconciliation, or public-sector data, budget 4 to 6 weeks. The ICP definition step alone typically takes a week if you’re doing it properly with cross-functional input.
Quarterly is the minimum. If you’re in a fast-moving market with high job turnover (tech, startups), monthly enrichment refreshes on your Tier A accounts are worth the cost. At minimum, re-validate email deliverability every 90 days and re-run phone enrichment every 6 months. Set calendar reminders. Nobody remembers to do this without a system.
AI helps at specific stages but doesn’t replace the process. LLMs are excellent for fuzzy deduplication, classifying companies into ICP tiers, and standardizing messy company names. AI-powered enrichment tools like Clay and FullEnrich use multiple models to find contact data. But ICP definition, source selection, and validation still require human judgment. Companies that fully automate TAM building end up with large databases of low quality.