Drupalwoo ☉

Creating a custom Drupal migrate source plugin for XML feeds

Drupal’s Migrate API handles structured imports well when the source matches an existing plugin. XML feeds often need a different approach: records may sit several levels deep, namespaces can obscure familiar elements, and identifiers may be buried in attributes rather than ordinary tags. A custom source plugin gives the migration a predictable interface without forcing feed-specific logic into every migration definition.

This pattern suits product catalogues, council notices, event listings, property feeds, membership systems, and legacy publishing platforms. It is useful when a JSON endpoint is unavailable or when an XML document contains richer metadata than the provider’s alternative format.

The example below targets modern Drupal development with PHP, Drupal 10 or 11, and a custom module. The same design works for a feed hosted on an Australian NBN connection, a vendor endpoint in Sydney, or a scheduled export produced by a Melbourne-based business.

A source plugin should do one main job: retrieve records and expose consistent source fields. Entity creation, text formatting, media handling, and deduplication belong in the migration definition or process plugins. Keeping those responsibilities separate makes future feed changes easier to manage.

Decide what the source plugin should own

A Migrate source plugin is the bridge between an external data set and Drupal’s migration pipeline. It tells Migrate which rows exist, which values each row contains, and which source property uniquely identifies a row. The destination might be nodes, taxonomy terms, users, media entities, or custom content entities.

For an XML feed, the plugin commonly handles the HTTP request, XML parsing, field normalisation, and source IDs. It should not decide whether a record becomes an article or a product. That decision belongs in YAML, where site builders can review it without changing PHP.

Consider an event feed containing this structure:

<event>
  <event_id>bris-4821</event_id>
  <title>Community garden workshop</title>
  <description><![CDATA[Learn seasonal planting techniques.]]></description>
  <start>2025-08-16T10:00:00+10:00</start>
  <location>
    <suburb>West End</suburb>
    <state>QLD</state>
  </location>
</event>

The source row could expose id, title, description, start, suburb, and state. Drupal then receives flat, useful values even though the original XML is nested. If the feed later changes from <event_id> to an id attribute, only the source plugin needs updating.

Stable identifiers are especially important. Do not use the row position, title, or publication date as the source ID. A provider may reorder events, change spelling, or publish two records with the same title. A durable external key such as bris-4821 allows Migrate to recognise an existing record during repeated imports.

Build the custom XML source plugin

Create a custom module such as feed_import, then add src/Plugin/migrate/source/XmlFeed.php. The source plugin can use the Migrate source base class and return an ArrayIterator containing normalised rows.

<?php

namespace Drupal\feed_import\Plugin\migrate\source;

use Drupal\migrate\Plugin\migrate\source\SourcePluginBase;

/**
 * Reads event records from an XML feed.
 *
 * @MigrateSource(
 *   id = "xml_feed"
 * )
 */
class XmlFeed extends SourcePluginBase {

  /**
   * {@inheritdoc}
   */
  public function fields() {
    return [
      'id' => $this->t('External event identifier'),
      'title' => $this->t('Event title'),
      'description' => $this->t('Event description'),
      'start' => $this->t('Start date and time'),
      'suburb' => $this->t('Australian suburb'),
      'state' => $this->t('Australian state or territory'),
    ];
  }

  /**
   * {@inheritdoc}
   */
  protected function initializeIterator() {
    $url = $this->configuration['url'];
    $response = \Drupal::httpClient()->get($url, [
      'timeout' => 30,
      'headers' => ['Accept' => 'application/xml'],
    ]);

    $xml = new \SimpleXMLElement((string) $response->getBody());
    $rows = [];

    foreach ($xml->event as $event) {
      $rows[] = [
        'id' => (string) $event->event_id,
        'title' => (string) $event->title,
        'description' => (string) $event->description,
        'start' => (string) $event->start,
        'suburb' => (string) $event->location->suburb,
        'state' => (string) $event->location->state,
      ];
    }

    return new \ArrayIterator($rows);
  }

  /**
   * {@inheritdoc}
   */
  public function getIds() {
    return [
      'id' => [
        'type' => 'string',
      ],
    ];
  }

}

The annotation is supported by established Drupal 10 versions. If the project uses a Drupal release and coding standard that prefers PHP attributes, use the corresponding MigrateSource attribute format documented for that version. The important part is that the plugin ID in PHP matches the plugin value in the migration YAML.

For production code, inject an HTTP client rather than calling the service locator directly. A plugin factory or a service-aware base class can provide Guzzle cleanly, making the class easier to test. Also catch request and parsing exceptions so a temporary vendor outage produces a useful migration error instead of an opaque PHP warning.

The example loads the entire response into memory. That is fine for a small event feed, such as a weekly export from a local council. It is a poor choice for a large catalogue. For large documents, use XMLReader to stream one record at a time, or ask the provider for pagination, date filters, or a compressed batch export.

Map XML values into a Drupal migration

The migration definition declares how the source rows become Drupal entities. Put the YAML in config/install for a deployable migration, or in a custom migrations directory managed by the project’s chosen workflow.

id: events_from_xml
label: Import events from XML
migration_group: feed_import

source:
  plugin: xml_feed
  url: 'https://example.org/events.xml'

process:
  title: title
  body/value: description
  body/format:
    plugin: default_value
    default_value: basic_html
  field_start:
    plugin: format_date
    source: start
    from_format: 'Y-m-d\TH:i:sP'
    to_format: 'Y-m-d\TH:i:s'
  field_suburb: suburb
  field_state: state

destination:
  plugin: 'entity:node'
  default_bundle: event

migration_dependencies: {}

The format_date process plugin converts the ISO-8601 value into the format expected by the destination field. Preserve the offset supplied by the feed whenever possible. Australian imports can cross AEST and AEDT, so blindly treating every timestamp as UTC may move a Brisbane event to the wrong displayed time or shift a Melbourne event around daylight-saving changes.

Add the source ID to the migration’s ids configuration when the migration needs explicit control over its high-water behaviour:

source:
  plugin: xml_feed
  url: 'https://example.org/events.xml'
  ids:
    id:
      type: string

For a custom source plugin, getIds() already defines the key, so the extra YAML is often unnecessary. The key requirement is that the source plugin returns the same ID on every run. If an upstream system supplies a numeric ID, cast it consistently to either a string or an integer.

Feeds frequently use XML namespaces. A document from a government or industry provider may contain elements such as event:title rather than title. Register the namespace before using XPath:

$namespaces = $xml->getDocNamespaces(TRUE);
$xml->registerXPathNamespace('event', $namespaces['event']);

foreach ($xml->xpath('//event:event') as $event) {
  // Normalise the namespaced record here.
}

CDATA sections are generally returned as strings by SimpleXML, but HTML inside descriptions still needs care. Decide whether the destination should receive filtered HTML, plain text, or a processed field. Drupal’s text format and text filtering should remain the final security boundary; never trust markup merely because it arrived through an XML feed.

If imported content needs interactive filtering or dependent fields, keep the migration focused on data transfer and add browser behaviour separately. For example, Drupal editors working with administrative forms may benefit from these AJAX form tips, but JavaScript should not be used to compensate for missing or inconsistent source data.

Handle errors, updates, and Australian operating conditions

A reliable importer validates the response before parsing it. Check the HTTP status, content type where practical, and whether the XML has the expected root element. A successful HTTP response containing an HTML maintenance page should fail clearly rather than create an empty import.

Use logging that identifies the migration, endpoint, and external record ID. That matters when a national organisation publishes feeds from several systems, or when a provider’s support team refers to an internal identifier that differs from Drupal’s node ID. Avoid logging full descriptions or personal information if the feed contains member, customer, or booking data.

Australian schedules deserve deliberate treatment. A feed generated in Perth may be timestamped in AWST, a Queensland source normally remains on AEST, and Sydney or Melbourne switches between AEST and AEDT. Store dates in a field configuration that matches Drupal’s timezone expectations, and test records around October and April rather than assuming every state follows the same clock rules.

The same practical discipline applies to connectivity. A migration run on a hosting account in Brisbane may reach a vendor endpoint reliably, while a local development machine on an office NBN service may encounter a firewall, proxy, or certificate issue. For repeatable deployments, prefer a controlled staging environment, set explicit timeouts, and avoid depending on a developer laptop for scheduled imports.

A migration should be rerunnable. Use --update when changing process logic, retain the source ID mapping, and decide how removed feed records should be treated. Some sites unpublish missing events; others retain historical records for reporting. That policy cannot be inferred safely from the XML response alone.

Checks that prevent avoidable import failures

Before running a live migration, confirm the source contract and the destination assumptions.

Pre-deployment checks

After a successful test import, monitor the operational behaviour rather than relying on one clean run.

Production safeguards

The word “arvo” may be common in an Australian editorial team, but a scheduler still needs an unambiguous timezone and a documented run time. A migration scheduled for 4:00 pm in Adelaide is not equivalent to 4:00 pm in Darwin, particularly when a central operations team manages sites across several states.

Choose the right parsing and delivery approach

SimpleXML is readable and productive for small and medium feeds. It gives developers direct access to nested elements and works well when the document structure is stable. Its main limitation is memory use, since the document is loaded before rows are yielded.

XMLReader is a better fit for very large exports because it streams through the document. It requires more code: the plugin must detect each target element, extract its XML fragment, and convert that fragment into a row. That extra complexity is worthwhile for large property, inventory, or archival feeds.

An HTTP source plugin is appropriate when Drupal should retrieve current data during each migration. A file-based source can be safer when a vendor places signed exports in object storage or when operations require an immutable copy of every import. In either case, keep acquisition and normalisation predictable.

Approach Strength Limitation Suitable use
SimpleXML Short, clear implementation Loads the full document into memory Small event or news feeds
XMLReader Handles large XML streams efficiently More complex row extraction Large catalogues and archives
HTTP retrieval in plugin Always reads the provider’s current feed Sensitive to network and vendor outages Scheduled synchronisation
Downloaded XML file Repeatable and auditable imports Requires file delivery and retention Regulated or high-volume workflows
Existing generic source plugin Less custom PHP to maintain May not understand unusual nesting or namespaces Conventional XML structures

Clear boundaries make the plugin maintainable. The source class should expose clean values, YAML should map those values, and destination-specific logic should use process plugins. With that separation, a feed from a Sydney events provider or a regional council system can change its transport details without forcing a rewrite of Drupal’s content model.