Skip to content
BlackBadger

Blog / Custom software

How do you migrate ten years of data out of a SaaS platform?

/ 9 min read

person using MacBook pro

TL;DR

Migrating a decade of data from a SaaS platform requires careful planning across data extraction, strategic decision-making about what to move, complex mapping work, and realistic expectations about what the process can accomplish.

Key takeaways

  • Run an actual test export before planning anything to discover what data is truly accessible, as file attachments, comment threads, and custom fields often export in formats that are incompatible or missing entirely.
  • Categorize your data into active data that moves fully with relationships intact, reference data that can move in simplified form, and dead data that should be archived or eliminated to dramatically reduce migration complexity and cost.
  • Build a translation table on day one that maps old system record IDs to new system IDs before import, and maintain it permanently because every linked relationship depends on this mapping to avoid broken connections.
  • Deduplication must happen in a staging database with human review of ambiguous matches, since automated matching alone will incorrectly merge different customers and create data problems that surface months later.
  • Check your contract for data retrieval clauses and API rate limits before giving notice to your vendor, as some platforms cut API access immediately after subscription lapses and slow rate limits can cause three-day data pulls that jeopardize your cutover timeline.
  • Ask your team specific questions about which records they actually use rather than whether they need history, because answers about warranty claims or occasional lookups reveal the minimum data required for the new system.

Ten years of data isn't one thing. It's customer records, sure. But it's also file attachments, comment threads, status history, custom fields somebody added in 2019 and never documented, and the audit trail your auditor asks for every spring. Getting it out of a SaaS platform is doable. It's also the part of a replacement project that people underestimate by the widest margin. Here's how the work actually goes, what breaks, and how to decide what's worth moving at all.

First, find out what you can actually get out

Before you plan anything, run a test export. Not a plan for one. An actual export, this week, of whatever your platform gives you. You'll learn more in two hours than in two weeks of vendor calls.

Most platforms hand you clean CSV files for the core tables. Contacts, companies, projects, invoices. That part usually works. The trouble starts one layer down. File attachments often come out as a separate download with cryptic filenames and no obvious link back to the record they belonged to. Comment threads and activity logs are frequently missing from the standard export entirely. Custom fields sometimes export with internal names like cf_00423 instead of "Preferred Install Window."

The API is usually the better path for anything complicated. An API is just a way for one system to request data from another, record by record, in a structured format. Most SaaS vendors offer one. Read the rate limits before you get excited. A limit of 100 requests per minute sounds fine until you calculate that 400,000 records at one record per request takes about three days of continuous pulling. That's fine if you plan for it. It's a disaster if you discover it the weekend of your cutover.

Check your contract for the data retrieval clause before you give notice. Some vendors cut API access the day your subscription lapses, and some only keep your data available for 30 days after that. Pull everything while you're still a paying customer in good standing.

Decide what actually needs to move

Not all ten years belong in your new system. This is the single biggest lever you have on cost and risk, and most teams skip the conversation.

Split your data into three buckets. Active data is anything your team touches now or will touch soon. Open jobs, current customers, unpaid invoices, live inventory. That moves, fully, with relationships intact. Reference data is stuff you look up occasionally but never edit. Closed projects from 2016. Former employees. Old quotes. That can move in a simplified form, or live in a read-only archive. Dead data is everything else. Test records, duplicate contacts, the imported list from a trade show nobody followed up on.

Take a distributor that comes in convinced it needs every historical order line moved into the new system. Ask who looks at orders older than three years, and the answer is usually one person, twice a year, for warranty claims. Move three years of full order detail and put the older records into a searchable archive table with the fields that warranty lookups actually use. The migration gets dramatically simpler, and nobody misses the rest.

Ask your team a specific question. Not "do you need the history?" Everyone says yes to that. Ask "when did you last open a record older than two years, and what were you looking for?" The answers are usually narrow and specific. Build for those answers.

The mapping work is the real project

Moving bytes is easy. Deciding what a byte means is hard.

Every migration turns into a series of small judgment calls. Your old system has a status field with 14 options, six of which are variations on "waiting." Your new system needs five. Which old status maps to which new one? Your contact records have a "Notes" field that people have been dumping phone numbers, PO numbers, and complaints into for a decade. Do you parse it, or import it as one blob and move on?

Then there are the identifiers. Every record in your old system has an ID, an internal number the platform uses to link things together. A project points at a customer by ID. An invoice points at a project by ID. If you import customers first and let the new system assign fresh IDs, every one of those links breaks unless you keep a translation table that says old ID 88231 is now new ID 4417. Build that table on day one. Keep it forever. You'll want it six months later when someone asks why a 2018 invoice is attached to the wrong job.

Duplicates are the other guaranteed headache. Ten years of manual entry means the same customer exists three times with slightly different spellings. Migration is your one clean chance to fix that. Do the deduplication in a staging area before the import, with a human reviewing the ambiguous matches. Automated matching on company name alone will merge two genuinely different customers and you won't find out until someone gets the wrong invoice.

Migrate to a staging database first, never straight into production. A staging database is just a private copy of the new system that only your project team can see. Load into it, break it, fix the script, wipe it, load again. Ten dry runs cost you nothing. One bad production import costs you a week of manual cleanup.

How the cutover actually runs

The pattern that works is boring and repeatable. Do a full test migration early, while you're still building. Load everything into staging. Then sit real users in front of it and ask them to find records they know well. Their own accounts. Last month's biggest job. They'll spot problems no reconciliation report catches, because they know what the data is supposed to say.

Then you run the real thing in two passes. The bulk load happens ahead of go-live, usually a week or two out. That's the ten years of history, which isn't changing anymore. The delta load happens at cutover, and only covers records that changed since the bulk load. That's a small, fast job you can do in an evening instead of a lost weekend.

Reconcile with numbers, not vibes. Count records in each table on both sides. Sum the dollar totals on invoices. Compare attachment file counts. If your old system says 41,208 invoices totaling a certain amount, your new one should match exactly. When it doesn't, the gap tells you where to look. Usually it's records with a null value in a required field, or dates in a format the import script didn't recognize.

Keep the old system in read-only mode for a while after cutover. Most vendors will let you drop to a single seat or a low tier. That's your safety net for the questions that surface in month two. Don't let anyone keep working in it though. Two live systems means two versions of the truth.

Attachments are almost always the ugliest part. A ten year old file store often holds tens of thousands of documents with duplicate names, broken links, and files uploaded by people who left in 2017. Budget real time for this. It's a scripted job that matches files to records using the export metadata, not a drag and drop.

What migration can't fix, and other honest caveats

A few things people expect from a migration that they shouldn't.

It won't clean bad data by itself. If your customer records are a mess in the old system, they'll be a mess in the new one. You can clean during migration, but that's human work and it takes hours. Decide up front how much you're willing to invest, and accept that the rest comes over as-is.

Some history is genuinely unrecoverable. Deleted records past the retention window are gone. Certain platforms don't expose the full edit history through any export or API, so "who changed this field and when" for prior years may simply not exist outside the vendor's own interface. If that history is a compliance requirement, find out now, not in month three. Sometimes the honest answer is that you keep a frozen archive of the old system for the retention period and start the new audit trail fresh.

Integrations don't come along for the ride. If your SaaS platform pushes data to accounting, or receives orders from a portal, every one of those connections gets rebuilt. That's separate work with its own testing, and it's often larger than the data migration itself.

Custom software also isn't the answer for everything. If the platform you're leaving is a commodity, like email or accounting, replacing it with a build is a bad trade. The case for owning your system is strongest where the software encodes how your business specifically works, and where you're paying per seat for a tool that still doesn't quite fit.

The arithmetic on leaving

Do this math with your own numbers. Take your per-seat monthly cost. Multiply by your seat count. Multiply by 12. That's your annual run rate. Now multiply by five, and add whatever your vendor's typical annual increase has been, compounded. That's your five year commitment if you renew.

Against that, put the one-time cost of a build you own, plus hosting and ongoing maintenance. Hosting for a system serving a few hundred users is a modest monthly bill. Maintenance is real and you should budget for it, but it doesn't scale with headcount. That's the structural difference. SaaS costs grow every time you hire. An owned system doesn't.

Add the migration cost to the custom side of the ledger, honestly. It's a real line item and pretending otherwise makes the comparison useless. Also add the switching cost you'll pay again in five years if you move from one SaaS platform to another instead. That cost is not zero, and people forget it.

If your seat count is growing and your platform fee is climbing faster than your revenue, the arithmetic usually favors owning. If you have 12 users on a cheap tool that does the job, it doesn't. Run the numbers before you run the project.

Where to start this week

Pull a test export. Count your records. Ask three people on your team what history they've opened in the past year. Read your vendor contract's data section. Those four things take a day and they'll tell you whether this is a straightforward job or a complicated one.

If you want a second set of eyes on what your data looks like and what it would take to move it, we do this work for businesses across Pinellas and Hillsborough counties. Book a free strategy session at /contact and bring your export file.

Questions

Frequently asked questions

What should I do before planning a data migration from a SaaS platform?

Run an actual test export this week, not just a plan for one. This will reveal more about what data you can retrieve in two hours than in weeks of vendor calls. Most platforms provide clean CSV files for core tables like contacts and companies, but complications arise with file attachments, comment threads, and custom fields.

What are the three categories for deciding which data to migrate?

Active data includes records your team touches now or will soon, like open jobs and current customers, which should move fully. Reference data is occasionally looked up but never edited, like closed projects, and can move in simplified form or to a read-only archive. Dead data is everything else, like test records and duplicates, which typically doesn't need to move.

Why is mapping data the real project in a migration?

Mapping determines what each data element means between systems. You must decide how old status values convert to new ones, handle messy fields like decades of notes, and manage identifier translation tables so relationships don't break. Small judgment calls add up to significant complexity that moving bytes alone cannot solve.

What is a staging database and why should I use it?

A staging database is a private copy of your new system that only your project team can access. Load, test, and break the migration here repeatedly before going live. Ten dry runs cost nothing, but one bad production import costs a week of manual cleanup and operational disruption.

How should a cutover migration actually be executed?

Run the migration in two passes. The bulk load happens before go-live and covers ten years of static history. The delta load happens at cutover and only covers records changed since the bulk load, completing in an evening instead of a weekend. Test migrations with real users first to catch problems reconciliation reports miss.

What contract clause should I check before starting a migration?

Check the data retrieval clause in your vendor contract before giving notice. Some vendors cut API access the day your subscription ends, while others keep data available for only 30 days afterward. Pull everything while you're still a paying customer in good standing to avoid losing access.

How do I handle duplicate records during migration?

Perform deduplication in a staging area before importing, with a human reviewing ambiguous matches. Automated matching on single fields like company name can incorrectly merge genuinely different customers. Manual review prevents mismatches that won't surface until customers receive wrong invoices.

Read this next

What happens to your data when you stop paying a SaaS vendor

Cancel a SaaS subscription and your data doesn't vanish, it gets locked behind a login. Here's what an export really gives you, what never comes out, and when owning the system is cheaper.

If this is your situation

Talk it through on a call

Thirty minutes over video with the person who would run the build. You talk, we ask questions, and you leave with a plain read on whether custom software makes sense for you. If the answer is that your current tool is fine and badly configured, you will hear that.

Book a strategy call

Who wrote this

David Verneuille

Founder, Black Badger Software Solutions

David runs Black Badger from Clearwater, Florida. He has spent years implementing the platforms this site writes about, which is why the writing here is specific about where they fit and where they stop fitting.

About Black Badger

Drafted with AI assistance, then checked, edited, and approved by a person before publishing. The facts, opinions, and recommendations are ours.

Related

More on custom software

Want this looked at properly?

An article can only go so far without knowing your systems. Book a strategy call, walk us through what is breaking, and you will get a straight read on whether custom software is the right answer.