Migrating WordPress Content Into Drupal With a Custom Migration Pipeline
Migrating a publishing site from WordPress to Drupal looks simple on paper. Export the content, point the new platform at the database, switch the domain. In practice every team discovers the same thing: the export is messy, the field models are different, and the moved content rarely behaves the way editors expect.
A custom migration solves most of those problems. Rather than relying on a contributed click-and-go module, you build the import around Drupal's Migrate API, the same framework core uses for its own upgrades. You get source plugins you control, process plugins that fit your content model, and YAML configuration you can version, review, and re-run as often as needed.
The work is mostly plumbing. WordPress stores posts, comments, and metadata in a handful of wp_* MySQL tables and offers a WXR (WordPress eXtended RSS) export. Drupal expects entities with fields, revisions, and translations. A custom migration sits between the two and translates one shape into the other without losing authors, images, or URL structure.
Why build it yourself instead of using a contributed module
The contributed migration ecosystem for WordPress is thin and uneven. Some modules handle only posts. Others break on custom post types or serialised meta. Even the better-maintained ones assume a default WordPress setup and struggle with ACF fields, multilingual content, or media stored off-server.
| Approach | Strengths | Weaknesses | Best fit |
|---|---|---|---|
| Direct SQL import via custom PHP | Full control over column mapping | No Migrate API integration, hard to re-run | Tiny personal sites |
| Contributed WordPress-to-Drupal module | Quick setup, opinionated defaults | Limited support for custom fields | Small standard blogs |
| Custom Migrate API pipeline | Repeatable, version-controlled, extensible | Higher upfront effort | Production sites with bespoke models |
| Hosted third-party migration service | Hands-off, handles complex schemas | Ongoing cost, vendor lock-in | Agencies with budget for tooling |
A custom Migrate API pipeline scales comfortably for an Australian publisher with regional offices, multilingual content, and a team that has been adding custom fields for years. The YAML files double as documentation, so the editorial lead in Brisbane can read them and see exactly how a legacy WordPress field ends up on a new Drupal node.
The biggest advantage is iteration. WordPress exports are rarely clean: a category might be stored as a slug in one row and as a name in another. With a custom source plugin, you write a row retrieval method that cleans up these inconsistencies before the rest of the pipeline ever sees them.
Source plugins that read WXR and WordPress databases
The source side of a migration is where most teams underestimate the work. WordPress can hand you content in two shapes: a live MySQL database, or a WXR XML file produced by the exporter. The database is faster to iterate against but couples your migration to live data; the WXR file is portable but must be regenerated whenever content changes.
For most projects, reading directly from the WordPress MySQL database is the right starting point. Add a database connection in settings.php, point your source plugin at wp_posts, wp_postmeta, wp_users, and wp_term_*, and let Migrate API pull rows in batches. A typical source plugin extends SqlBase and overrides the query, fields, and unique identifier. WordPress uses auto-increment IDs that change between exports, so use the post GUID or a hash of the title and publication date instead.
If the legacy site is on shared hosting and the database is unreachable, the WXR file is the fallback. WXR is essentially an RSS feed with extra namespaces for attachment metadata and custom fields. A small custom source plugin using XMLReader can stream through a several-hundred-megabyte export without exhausting PHP's memory limit on a standard Australian VPS.
Both approaches should expose the same fields downstream. Whether the row came from SQL or XML, the process plugins should see title, body, author, published_at, and a list of taxonomy terms. Hiding the difference at this layer keeps the rest of the configuration readable.
Process plugins for authors, taxonomy, and media
Process plugins are where the migration actually becomes yours. Drupal's Migrate API ships with built-in processors like skip_on_empty, concat, and entity_lookup, but WordPress content rarely fits those patterns.
Authors are usually the first headache. WordPress stores them in wp_users with a separate wp_usermeta table holding display names and roles. Drupal wants user entities with roles assigned through configuration. A common pattern is to look up the Drupal user by email and create one if missing, using a custom process plugin that wraps entity_lookup with a creation callback. That preserves attribution and avoids orphaned "anonymous" nodes.
Taxonomy mapping is where you earn the migration's keep. WordPress categories and tags live in a single wp_term_taxonomy table with a parent-child column. Drupal separates vocabularies from terms and expects hierarchical terms to reference parents by ID. A custom process plugin walks the term tree, creates each term in the right vocabulary, and rewrites the parent reference to the newly assigned Drupal term ID.
A short checklist that keeps process plugins honest:
- Validate every custom plugin with a unit test that hands it a known input.
- Log skipped rows with a reason rather than failing the import.
- Use
skip_on_emptyfor optional fields andrequiredfor fields the editor cannot save without. - Cache expensive entity lookups to keep the migration's runtime reasonable.
Preserving SEO with redirects and aliases
The most visible failure of a WordPress-to-Drupal migration is a sudden drop in search traffic. Google has spent years learning the old URLs, and silently 404-ing them after launch undoes that work overnight.
Drupal's redirect module accepts CSV imports of source and destination paths. Export the URL list from WordPress using a script that reads wp_posts.guid or builds paths from the slug and published date, then hand the CSV to the redirect importer. For larger sites, write a custom Migrate API destination plugin that creates a redirect entity per row instead.
Pathauto is the other half of the picture. Configure the new site's path patterns to produce URLs that closely match the old structure, then add a process step that writes the old WordPress path into the Drupal node's alias field when the patterns diverge. This keeps editorial workflows predictable while letting the redirect module cover every URL the old site served.
Do not forget xmlrpc.php and wp-login.php. Drupal does not ship with them, but WordPress clients and aggregators will keep hitting them for months. Serve them with a 410 Gone or block them at the web server. Both are quick wins for security and for keeping server logs clean.
Hosting, time zones, and compliance on Australian infrastructure
Once the pipeline runs cleanly on a developer laptop, the next questions are operational. Where will the Drupal site live, what time zone will it run in, and how does the migration interact with Australian privacy law?
Most Australian teams choose local providers such as PanAustralian CloudNine in Sydney, VentraIP in Melbourne, or Digital Pacific in Brisbane, often behind the NBN at the office and a CDN at the edge. Pick a host that lets you take a snapshot before each migration run. Migrate API's database rollbacks are useful, but a host-level snapshot is the safety net you want when touching a live WordPress database.
Set the Drupal site to Australia/Sydney (or Australia/Adelaide if you operate across ACDT) before importing any content. WordPress stores timestamps in UTC and displays them in the configured display time zone. Drupal does the same, but if the two sites are configured differently, scheduled posts and comment timestamps will drift by hours. Running drush config:set system.date_timezone.default Australia/Sydney is a one-line fix that prevents a confusing bug report later.
Privacy compliance matters because WordPress often collects more than editors realise. Commenter email addresses, contact-form submissions, and old WooCommerce orders sit in the WordPress database and would otherwise migrate into the new site. Under the Privacy Act 1988 and the Australian Privacy Principles (APPs), personal information about identifiable individuals should only be retained as long as it is needed. Strip commenter PII during the migration, replace real email addresses with anonymised placeholders, and document the retention decision in the runbook.
For agencies working with federal or NSW government clients, the Australian Cyber Security Centre's Essential Eight maturity model is the baseline. The migration script runs with database credentials that touch the entire legacy site, so use a dedicated, short-lived MySQL user, store the password in the project's secrets manager, and rotate it after the migration is complete.
A handful of small operational habits make the exercise smoother:
- Schedule the cutover to avoid AEST public holidays, when editorial traffic is low but on-call coverage is hardest.
- Run the migration against a copy of production first, then again the night before launch, then once more as a dry run an hour before the real cutover.
- Keep the legacy WordPress instance online behind a separate hostname for at least a month in case a stray URL needs to be reconstructed.
- Brief the editorial team in Melbourne or Perth a week ahead so they can flag any content the migration might miss.
A custom Drupal migration from WordPress is rarely a single afternoon's work, but it is also not the multi-week project it can feel like when you first open a WXR file. With a source plugin that speaks the legacy format, process plugins that fit your content model, and a redirect strategy that protects search rankings, the migration becomes a controlled, repeatable process rather than a leap of faith.