Advanced Drupal Search with Search API and Solr
Search functionality often decides whether visitors stay or bounce. On Australian government and enterprise sites — many of which run on Drupal, including portals under the data.gov.au ecosystem — users expect sub-second responses even when querying millions of records. Building that kind of performance with the default database backend is unrealistic, which is why most teams turn to the Search API module paired with Apache Solr.
The combination is flexible, scalable, and well-documented, but it requires careful configuration to deliver relevance that matches user intent. This walkthrough covers the architecture, installation, indexing strategies, and tuning techniques you need to ship a production-grade search experience on Drupal.
Why Solr Beats the Database Backend
Drupal's built-in search relies on SQL LIKE queries, which stumble quickly once content volumes grow. A site that indexes a few thousand nodes might cope, but anything beyond that — multilingual corpora, complex taxonomies, faceted product catalogues — grinds to a halt. Solr solves this with an inverted index, tokenisation, and a relevance model borrowed from Lucene that has been refined for two decades.
For Australian organisations, this matters in practical ways. A Sydney-based media publisher serving national audiences cannot afford search to lag during morning traffic peaks in AEST. A university in Melbourne indexing tens of thousands of thesis abstracts needs facets by faculty, year, and supervisor — features that SQL simply cannot deliver efficiently. Regional councils in places like Hobart or Cairns deal with growing document libraries that quickly outgrow database search, and they have no patience for queries that take five seconds to return.
Beyond raw speed, Solr offers linguistic tools that align with Australian English. Stemming, stop words, and synonym filters can be tuned so that a search for "colour" matches documents containing "colors", and vice versa. This kind of normalisation is fiddly to replicate with SQL and almost impossible to scale beyond a few thousand records.
Installing Solr and Wiring It to Drupal
Start by provisioning Solr. Many Australian Drupal hosts — including several with points of presence in Sydney and Melbourne — provide managed Solr services that handle replication, backups, and uptime. If you prefer self-hosting, Solr 9.x runs cleanly on Ubuntu or RHEL and pairs nicely with Drupal 10 and 11.
The Drupal side requires three contributed modules: Search API, Search API Solr, and optionally the Solr Admin module for a UI bridge. Composer handles installation cleanly with composer require drupal/search_api drupal/search_api_solr. Once enabled, configure a Solr server at /admin/config/search/search-api. Point it at your Solr core URL, choose the appropriate connector, and supply the core name. The Search API Solr module ships configuration files that match the schema Drupal expects — copy them into your Solr core's conf directory, restart Solr, and you are ready to index.
If you'd rather not maintain Solr infrastructure yourself, see about this site for managed options that absorb weeks of DevOps work. Just make sure the Solr version matches your Drupal core, because mismatched schemas are the most common cause of silent indexing failures that surface only when someone notices a missing result.
Index Design and Backend Trade-offs
A search index is a flattened representation of your content, tuned for retrieval rather than display. The art lies in deciding what to include. For a typical Drupal site, you want the title, body, tags, taxonomy terms, and any custom fields that carry meaningful text. Numeric fields like publication year or price should be typed as such, because Solr uses them differently from text and lets you filter and sort with native efficiency.
Processors transform data before it hits the index. The HTML filter strips tags so that markup does not pollute token streams. Ignore Case normalises capitalisation, and a Tokenize by Stopwords processor removes common function words that add noise. A useful exercise is to think about the kinds of queries real users run. A Brisbane legal publisher might need to search by jurisdiction (Queensland, NSW, federal), by document type, and by date. Each of these dimensions becomes a fulltext or string field in the index. For multilingual sites, the Language processor splits content per language so that stemming does not bleed across boundaries.
Here is a quick comparison of common backend choices:
| Feature | Database Backend | Solr Backend |
|---|---|---|
| Indexing speed | Slow on large corpora | Fast, optimised for batch loads |
| Full-text relevance | Basic LIKE matching | BM25 with custom scoring |
| Faceting | Manual SQL aggregation | Native, multi-select |
| Highlighting | Not supported | Native with configurable tags |
| Multilingual support | Limited | Strong, per-language analysis |
| Operational overhead | None | Cluster and schema management |
For projects that need facets, highlighting, typo tolerance, and reasonable performance past a few thousand items, Solr is the only sane choice. The operational cost is real, but it pays for itself the moment a user finds what they need without paging through dozens of irrelevant results.
Tuning Relevance and Designing the Search Page
Indexing is half the battle. The other half is shaping how queries resolve. Solr's qf parameter tells the engine which fields to search and with what boost. A common pattern is title^5 body^2 tags^3, which makes matches in the title weigh heavily, matches in tags slightly less, and body matches count but less decisively.
Synonyms are where relevance tuning becomes editorial. For an Australian news site, "Liberal Party", "Libs", and "Coalition partner" might all map to the same cluster. Solr exposes a managed synonyms file that you can edit without reindexing everything — a feature worth its weight in gold when editors want to chase a news cycle. Spell checking, "Did you mean" suggestions, and stemming exceptions all live in similar flat files that the Search API Solr module maps to Drupal configuration.
The user-facing search page deserves care too. Drupal's Views module integrates with Search API through a dedicated display plugin, which means you can build a search results view with exposed filters, facets, and pagination without writing custom code. A clean layout includes a results count, sortable columns, facet blocks in a region, and an empty-state message that suggests popular terms. For local government sites in places like the City of Adelaide or the City of Perth, an autocomplete box driven by Solr's suggester component makes a tangible difference to usability, especially on mobile devices where typing is slower.
Scaling, Maintenance, and External Sources
Solr is not a set-and-forget system. Index size grows, query patterns shift, and relevance models need periodic recalibration. The Solr Admin UI gives you a dashboard for cache hit rates, query timings, and index size. Pair it with application-level logging so you can spot slow queries and missing-result patterns. APM tools that integrate with Australian observability platforms can send alerts when latency creeps above a threshold.
Scaling options include vertical scaling (bigger Solr boxes), horizontal scaling (SolrCloud with sharded collections), and read replicas. SolrCloud is the right answer once you cross a few million documents or need high availability. It introduces a ZooKeeper ensemble for cluster coordination, which is one more thing to operate — another reason many Australian teams lean on managed services rather than running their own clusters.
Some projects also need to search content that lives outside Drupal — think archived PDFs, a separate WordPress blog, or even niche sites such as a photo gallery covering honey harvest traditions. Solr handles this gracefully through its DataImportHandler or direct HTTP push, with Drupal exposing the results via a custom datasource plugin. The trick is keeping a consistent document schema and indexing external content on a slower schedule, since latency tolerance is usually higher than for first-party content.
The reward for this discipline is a search experience that feels effortless to users and predictable to operators, exactly what Drupal sites need when content volume and audience expectations keep climbing.