Skip to content
Shops & CMS

Instant Search Suggestions: Speeding Up Your Shop Search

Shop search as a real-time path: debounce input, abort outdated requests, cache frequent prefixes and render a lean suggestion list.

15 min read INPE-CommerceCachingJavaScriptShopsuche

Anyone typing into a shop's search field usually already knows what they want. No other path through a shop shows purchase intent so clearly - and hardly any other responds so directly to every single key. Search suggestions have long been standard: 80 percent (Baymard Institute) of the benchmarked online shops offer them, and in the first benchmark in 2014 it was 72 percent (Baymard Institute). That makes it all the more noticeable when the list lags behind, flickers or shows suggestions for a term nobody is typing anymore. This article treats search as a real-time path from keystroke to painted suggestion list, measures it against the perception limits of 0.1 and 1.0 seconds (Nielsen Norman Group) and shows the matching fix for each segment: debouncing, aborting outdated requests, a response cache for frequent prefixes and a lean result list. For online shops and B2B portals, this is a fixed building block of e-commerce performance.

Key takeaways

  • Search suggestions are standard: 80 percent (Baymard Institute) of the benchmarked online shops offer them. Search is the path with the clearest purchase intent - every delay hits visitors who already know what they want.
  • Every keystroke in the search field counts as an interaction (web.dev) and therefore feeds into INP. An INP of 200 milliseconds (web.dev) or less counts as good - heavy input handlers and a long render phase of the list hit this value directly.
  • Perception provides the targets: up to 0.1 seconds (Nielsen Norman Group) a response feels instantaneous, up to 1.0 seconds (Nielsen Norman Group) the flow of thought stays uninterrupted. The typed character belongs below the first limit, the suggestion list below the second.
  • Debouncing and aborting belong together: the request only starts during a typing pause, and every new search aborts the one still running via AbortController (MDN). That way the list never shows suggestions for an outdated term.
  • Frequent prefixes repeat across all visitors. A response cache with a normalized key answers them without asking the database; sales channel, language and customer group belong in the key, session data does not.
  • Search field and results page are missing from typical Lighthouse runs because nobody types there. Measure in timespan mode with real input and in the field with RUM data filtered to the search field.

Why shop search carries the clearest purchase intent

A visitor who arrives via the navigation is browsing. A visitor who types into the search field is looking for something - a specific item, a brand, an order number from the last quote. In a B2B portal this is even more pronounced: buyers know their item numbers and expect the right line to appear in the list after a few characters. Search is therefore the path through the shop where intent and patience are furthest apart. Intent is high, willingness to wait is low. Anyone typing into the void here loses the thread before the shop has even had a chance to show what it carries.

Search suggestions are no longer an extra but an expectation. In its e-commerce benchmark, the Baymard Institute finds that 80 percent (Baymard Institute) of the shops studied offer search autocomplete; in the first benchmark in 2014 it was already 72 percent (Baymard Institute). Anyone who has experienced a search field with suggestions reads every pause as a disruption. At the same time, autocomplete is technically one of the most demanding parts of a shop: it combines input processing in the browser, a network call, a database or index lookup and the rendering of a list - and not once per page view, but potentially on every keystroke.

That is exactly why looking at the load time of the home page is not enough here. A shop search can sit on a page with a good Largest Contentful Paint and still feel sluggish, because its time is not spent while loading but during use. How this usage time shows up in the metrics is explained in our overview of the Core Web Vitals.

Search is usage time, not load time

Most performance work targets the initial build of the page. Search suggestions come afterwards: they are a chain of interactions whose cost repeats with every input. A shop can shine in the lab and still stall in the search field. That is why search is measured as a path of its own - with real input, on a realistic device and across the whole chain from keystroke to list.

The search path from keystroke to painted list

Between pressing a key and the finished suggestion list there are three segments, which are measured differently and fixed differently. The first runs on the browser's main thread: the input event is processed and the typed character appears in the field. This segment is captured by Interaction to Next Paint. web.dev explicitly lists pressing a key on either a physical or an onscreen keyboard among the interactions INP assesses, and names a value of 200 milliseconds (web.dev) or less as good responsiveness. Heavy input handlers, synchronous filtering of large lists in the browser or an elaborate repaint lengthen exactly this segment. How to lower INP in general is shown in our article on optimizing INP.

The second segment is the network path: call to the server, lookup in the index or database, response back. It does not appear in the INP of the keystroke, because the browser has long since painted the typed character before the response arrives. For the visitor it still matters - it decides when the suggestions appear. The third segment is rendering the list once the response is there. It runs on the main thread again, and if it takes long, it blocks the next input: the following keystroke has to wait until the list is finished, and this waiting time enters the next interaction as input delay.

The perception limits described by the Nielsen Norman Group work well as targets. Up to 0.1 seconds (Nielsen Norman Group) people perceive a response as instantaneous; no special feedback is needed, it is enough to display the result. Up to 1.0 seconds (Nielsen Norman Group) the flow of thought stays uninterrupted, even though the delay is noticed. For search this means: the typed character belongs below the first limit, the suggestion list for the current term below the second. Anything beyond that interrupts visitors in the middle of their input.

0.1 second is about the limit for having the user feel that the system is reacting instantaneously, meaning that no special feedback is necessary except to display the result.

Nielsen Norman Group, Response Times: The 3 Important Limits
SegmentWhat happens thereTypical bottleneckFix
Input on the main threadProcess event, paint characterHeavy input handler, filtering in the browserKeep the handler lean, defer work
Network and serverCall, index lookup, responseOne call per key, uncached lookupDebounce, abort, prefix cache
Rendering the listTurn hits into markup and paintLarge images, many nodes, layout shiftsShort list, fixed dimensions, reuse nodes
Next inputFollowing keystroke is processedList still blocks the main threadSplit rendering, give the browser room

The core in one sentence

The character in the field must appear at once, the list may wait briefly - but never so long that it shows a term nobody is typing anymore, and never so heavy that it holds up the next keystroke.

Debouncing: not every keystroke needs a request

The most common bottleneck is also the simplest: the search field fires off a request on every keystroke. Anyone typing a longer term creates as many requests as the term has letters - and almost all of them answer a question that nobody is asking anymore by the time the response arrives. Each of these requests occupies a connection, a worker on the server and, in the end, processing time in the browser when the response is handled. During peak loads in the busy season, this waste quickly becomes a bottleneck for the whole shop.

Debouncing solves this: the request only starts once the input has been idle for a brief moment. If the visitor types quickly, one request is created at the end of a typing pause instead of one per character. The pause is not guessed but tuned against your own measurements: long enough to bundle fast typing, short enough that pause, server response and rendering together stay below the limit of 1.0 seconds (Nielsen Norman Group). A minimum term length belongs with it. A single letter rarely yields useful suggestions, but it does yield an expensive lookup with a huge result set - in product ranges with short item numbers, on the other hand, the threshold can be lower than in a fashion shop.

What matters is what is not debounced: painting the typed character. The browser displays the character itself as soon as the main thread is free. Anyone who synchronously filters a list, sends analytics events or repaints half the search module in the input handler lengthens the processing time of every single key. The handler should only take the value and restart the timer - everything else runs later. How to split up unavoidable long tasks is shown in our article on scheduler.yield and long tasks.

search-field.js
const field = document.querySelector('#search');
const PAUSE_MS = 150;   // starting value, tune against your own field data
const MIN_CHARS = 2;    // starting value, check per product range
let timer;

// The handler only takes the value and restarts the timer.
field.addEventListener('input', () => {
  clearTimeout(timer);
  const term = field.value.trim();
  if (term.length < MIN_CHARS) {
    clearSuggestions();
    return;
  }
  timer = setTimeout(() => loadSuggestions(term), PAUSE_MS);
});

Debouncing does not replace a fast server

A typing pause reduces the number of requests, but not the duration of each one. If the server response is slow, the list arrives even later with debouncing than without, because the pause is added to the response time. Debouncing is therefore the first step, not the only one: only together with a fast, cached response do you get a list that stays below the one-second limit.

Aborting outdated requests

Even with debouncing, several requests are often in flight at once: the visitor types, hesitates briefly, types on. The responses do not necessarily come back in the order in which the requests were started. If the request for a short prefix takes longer - say because it searches more hits -, it arrives after the response for the longer term and overwrites the correct list with an outdated one. The visitor sees suggestions for a term they are no longer typing, and the browser has rendered an entire list for nothing.

The browser ships the solution: an AbortController can abort one or more running requests at any time (MDN). Every new search aborts the previous one before it starts itself. The aborted request no longer delivers a response that would have to be rendered. A simple check is also worthwhile: before rendering, the code verifies that the response still matches the current field content. Together, both prevent flickering between old and new suggestions and save the main thread exactly the rendering work that would otherwise hold up the next key.

suggestions.js
let pending;

async function loadSuggestions(term) {
  pending?.abort();                     // abort the outdated request
  pending = new AbortController();
  try {
    const response = await fetch(
      `/api/suggestions?q=${encodeURIComponent(term)}`,
      { signal: pending.signal }
    );
    const data = await response.json();
    // Only render if the term is still in the field
    if (term === field.value.trim()) renderList(data);
  } catch (error) {
    if (error.name !== 'AbortError') throw error;
  }
}

Aborting has one limit: it ends the request from the browser's point of view, not necessarily the work on the server. Whether a database lookup that has already started stops depends on the backend; often it runs to completion and its result is discarded. That is why the server segment remains a task of its own. Aborting saves rendering and bandwidth, debouncing saves requests, and only the cache saves processing time on the server.

The server side: frequent prefixes from the cache

Search terms are not evenly distributed. The beginnings of terms - the prefixes - necessarily repeat more often across all visitors than whole terms: every complete term starts with its prefixes, and many different terms share the same beginning. These are often the most expensive lookups, because short prefixes produce the largest result sets. And they are the easiest to cache, because their response is identical for all visitors in the same context.

A response cache for search suggestions needs a clean key. The search term is normalized first: trim spaces at the edges, unify upper and lower case, align Unicode spellings. Otherwise variants of the same prefix end up as separate cache entries and the hit rate drops. The key must include everything that actually changes the response - sales channel, language, currency and, in a B2B portal, the customer group if suggestions show prices or customer-specific assortments. Session identifiers or tracking parameters do not belong in it: they make every entry unique. How a single cookie wrecks the hit rate is described in our article on cookies, Vary and cache fragmentation.

For lookups that miss the cache, the lookup itself counts. A regular B-tree index only helps a pattern search if the pattern is anchored to the beginning of the string - that is with LIKE 'drill%', not with LIKE '%hammer' (PostgreSQL documentation). Without a usable index, the database is usually left scanning the table row by row. A prefix search or a dedicated search index with precomputed word beginnings answers the same question with a fraction of the work. Which lookups cost the most time in a shop is shown in the article on database query optimization; how much of it falls on a single search call is made visible by a response header that we describe in the article on Server-Timing.

Prefix cache with a short lifetime

Frequent prefixes are cached as a finished response. A short lifetime keeps stock and prices current without every input reaching the database.

Normalized key

Edge spaces, upper and lower case and Unicode variants are unified before the lookup. Channel, language and customer group belong in the key, session data does not.

Lean response

The endpoint only delivers what the list shows: name, URL, small preview image and, if needed, the price. No complete product objects and no fragments with embedded scripts.

Warm the cache before the visitor types

Frequent prefixes can be determined from your own search logs and refilled specifically after every import or price run. Then even the first request after an update hits a warm cache. If the search endpoint sits on its own domain, the browser can open the connection as soon as the search field receives focus. Which cache layer suits which purpose is covered by our caching strategies.

Rendering the suggestion list lean

Once the response is there, the third segment begins. The suggestion list is a small piece of interface that is rebuilt again and again in quick succession - and that is why every element in it counts. A list with large product images, detailed price blocks, rating stars and scripts of its own per entry has to be rebuilt, laid out and painted on every update. If that runs in a single long task, the next keystroke waits until it is done.

The most effective measures are unspectacular. Limit the number of entries: a short list is painted faster and easier to scan than a long one; anyone who wants to see more ends up on the results page anyway. Deliver preview images small - at the size at which they are displayed, in a modern format and with fixed dimensions, so the list does not jump when images load. Reuse existing nodes instead of discarding and recreating the whole list on every response. And load everything that does not need to be visible immediately, such as stock levels or tiered prices, only on the product page. Why every additional node counts is described in our article on DOM size.

Keyboard operation is part of a lean list. Arrow keys, Enter and Escape are keystrokes that INP assesses as well. Anyone who rerenders the whole list when the highlight moves, instead of just moving the highlight, lengthens every one of these interactions. The right markup - an input with the role combobox, a list with the role listbox and aria-activedescendant for the highlighted entry - helps screen readers and costs no extra render time. If a lot has to be painted after all, it helps to split rendering into smaller tasks so that a waiting input can be processed in between.

Few entries

A short, well-sorted list is painted faster and grasped faster. More hits belong on the results page, not in the dropdown.

Small images, fixed dimensions

Preview images at display size and in a modern format, with width and height in the markup. That way the list does not jump when the images arrive.

Reuse nodes

Update entries instead of discarding them: swap text and image, keep the structure. This saves style and layout calculation on every response.

A rule of thumb for the list

The suggestion list is not a small catalog but a signpost: paint as little as possible and lead the visitor to the right page as quickly as possible.

Measuring search pages that are missing from Lighthouse runs

Typical Lighthouse runs load a URL, wait until the page is quiet and assess how it was built. Nobody types. Autocomplete is never triggered this way, and the search results page is usually not even on the list of tested URLs, because it only comes into being through input. A shop can therefore look good in every report while its entry point with the clearest purchase intent remains unmeasured. How to put Lighthouse results in context is shown in the article on the Lighthouse report.

Search needs three perspectives. In the lab, Lighthouse in timespan mode analyzes an arbitrary period of time that typically contains user interactions - such as typing into the search field and submitting the search - and measures layout shifts and JavaScript execution time within it (Lighthouse documentation). The Performance panel of the developer tools shows every single key with input delay, processing time and presentation delay; with a throttled CPU, this comes closer to a mid-range smartphone than the work computer does. In the field, RUM data provides the real INP, and if every interaction records the element that triggered it, INP can be filtered to the search field. What field data can do and where CrUX stops is compared in our article on RUM and CrUX.

The network path is measured separately: how many requests does a typical search term create, how long does each take, how many are aborted, and which prefixes miss the cache? The browser's Resource Timing data and a Server-Timing header on the search endpoint answer this without additional tools. The results page itself belongs in every measurement series with a typical search URL - as a template of its own, just like category and product pages. Why the home page alone is never enough is explained in the article template performance beyond the home page.

  1. Define typical search terms: an item number, a brand name, a general term
  2. In timespan mode, type each term at normal speed, select a suggestion and open the results page
  3. Repeat with a throttled CPU and record input delay, processing and presentation for each key
  4. In the Network panel, count the requests per term and note duration and aborted requests
  5. Read Server-Timing on the search endpoint: cache hit or database lookup?
  6. In the field, filter INP by target element and evaluate the search field separately from the rest of the page
  7. Add the results page with a real search URL to the regular measurement series

The implementation path in practice

A fast shop search does not come from a single trick but from tuning all three segments together. The reliable path starts with measurement: where does the path lose time - in the input handler, on the server or while rendering the list? Only then is it decided whether to reduce requests, cache responses or slim down the list first. In Shopware shops, custom storefronts and B2B portals the levers are the same, only their order differs.

  • Measure the search path: input on the main thread, network path and list rendering separately
  • Keep the input handler lean - only take the value and restart the timer
  • Debounce requests and set a minimum length for search terms
  • Abort every outdated request via AbortController and discard stale responses
  • Cache frequent prefixes with a normalized key - channel, language and customer group in the key
  • Check the lookup itself: prefix search or a dedicated index instead of a leading wildcard
  • Keep the list short, images small with fixed dimensions, reuse nodes
  • Include search field and results page in lab and field measurement

We implement this chain as part of Shopware performance and in custom shop systems - from the input handler through the search endpoint to the list. Where the server response is the bottleneck, the work starts with server optimization; for portals with customer groups and tiered prices, the specifics of PageSpeed optimization for B2B portals apply. And which segment of your search actually slows things down is shown by a structured performance analysis before a single line is changed.

Sources and Studies

This article is based on data from: Baymard Institute (Autocomplete Design), Nielsen Norman Group (Response Times: The 3 Important Limits), web.dev (Interaction to Next Paint), MDN Web Docs (AbortController), PostgreSQL documentation (Index Types) and Lighthouse documentation (User Flows in Lighthouse). All figures cited were verified at the time of publication.

Related Articles

Front-end optimisation

Frontend Memory Leaks: When Long Sessions Slow Down

The interface starts fast and turns sluggish after twenty minutes: how front-end memory leaks arise, how to measure them over time and how to get rid of them.

12 min read
Core Web Vitals & measurement

Web Workers: Offload the Main Thread to Improve INP

Offloading compute-heavy JavaScript to a web worker keeps the main thread free, lowers Total Blocking Time and brings INP under the 200 ms threshold.

12 min read
Core Web Vitals & measurement

Break Up Long Tasks: Lower INP with scheduler.yield

Break long JavaScript tasks over 50 ms into chunks with scheduler.yield(), free the main thread between the chunks and keep the INP steadily under 200 ms.

12 min read