Elasticsearch in Practice: Mapping Design, Analyzers, Shard Sizing and Relevance Tuning

Key takeaways

Elasticsearch enables sub-second full-text search across millions of documents, but its query-time model differs fundamentally from SQL. This guide covers index design, mapping immutability and reindexing, the analyzer mismatch pitfall, shard sizing trade-offs, relevance scoring, and practical Node.js integration patterns.

What Elasticsearch Does

Elasticsearch is a distributed search and analytics engine built on Apache Lucene:

PostgreSQL query:  SELECT * FROM products WHERE name LIKE '%wireless headphones%'
  → Scans every row, case-sensitive, no ranking

Elasticsearch:    search "wireless headphones" in products
  → Inverted index lookup (O(1)), relevance scoring, fuzzy matching,
    handles "headphone" → "headphones", returns best matches first

The reason this distinction matters in practice is that a LIKE '%...%' query in a relational database cannot use a B-tree index at all — the leading wildcard forces a full table or sequential scan, and it has no concept of ranking, so every match is equally “correct” regardless of relevance. Elasticsearch instead builds an inverted index: for every distinct token, it stores the list of documents that contain it, plus enough statistics (term frequency, document frequency) to compute a relevance score at query time. That data structure is what makes sub-second search over tens of millions of documents possible, and it’s also why Elasticsearch is a fundamentally different tool from a relational database rather than a faster version of the same thing — you don’t get transactions or joins, but you get ranking, fuzzy matching, and language-aware tokenization for free.

Use cases:

  • Full-text product/content search
  • Log analytics (ELK Stack: Elasticsearch + Logstash + Kibana)
  • Autocomplete and suggestions
  • Faceted search (filter by category, price range, brand)
  • Real-time dashboards from event streams
flowchart LR
    A["Document write"] --> B["Analyzer<br/>tokenize + normalize"]
    B --> C["Inverted index<br/>term → doc list"]
    D["Search query"] --> E["Same analyzer<br/>(must match!)"]
    E --> F["Term lookup"]
    F --> C
    C --> G["BM25 scoring<br/>+ ranked results"]

The diagram above is the single most important mental model in this guide: the same analyzer has to run on both sides — once when a document is indexed, and once when a query is parsed — or the terms being compared won’t match even though they look identical to a human reading them. The rest of this guide comes back to that point repeatedly because it is the single most common production bug in Elasticsearch deployments.


Setup

# Docker — easiest for development
docker run -d \
  --name elasticsearch \
  -p 9200:9200 \
  -e "discovery.type=single-node" \
  -e "xpack.security.enabled=false" \
  elasticsearch:8.12.0

# Verify
curl http://localhost:9200
# { "name": "...", "version": { "number": "8.12.0", ... } }
# Node.js client
npm install @elastic/elasticsearch
// lib/elasticsearch.ts
import { Client } from '@elastic/elasticsearch';

export const es = new Client({
  node: process.env.ELASTICSEARCH_URL || 'http://localhost:9200',
});

Index and Mappings

Define the schema before indexing documents:

// Create index with explicit mapping
await es.indices.create({
  index: 'products',
  body: {
    settings: {
      number_of_shards: 1,
      number_of_replicas: 1,
      analysis: {
        analyzer: {
          // Custom analyzer for product search
          product_search: {
            type: 'custom',
            tokenizer: 'standard',
            filter: ['lowercase', 'asciifolding', 'stop'],
          },
        },
      },
    },
    mappings: {
      properties: {
        name: {
          type: 'text',
          analyzer: 'product_search',
          fields: {
            keyword: { type: 'keyword' },       // Exact matching and sorting
            suggest: { type: 'completion' },    // Autocomplete
          },
        },
        description: {
          type: 'text',
          analyzer: 'product_search',
        },
        price: { type: 'float' },
        category: { type: 'keyword' },          // Exact match, aggregations
        brand: { type: 'keyword' },
        rating: { type: 'float' },
        inStock: { type: 'boolean' },
        tags: { type: 'keyword' },
        createdAt: { type: 'date' },
      },
    },
  },
});

Field type guide:

  • text — analyzed, for full-text search (tokenized, lowercased)
  • keyword — not analyzed, for exact matching, sorting, aggregations
  • float, integer, long — numeric fields
  • date — date/datetime fields
  • boolean — true/false
  • completion — for autocomplete/suggestions

Mapping Is (Mostly) Permanent — Plan Before You Index

This is the part of Elasticsearch that catches people coming from a relational-database background off guard, and it deserves more attention than most tutorials give it. In PostgreSQL, ALTER TABLE products ALTER COLUMN price TYPE numeric is a routine migration. In Elasticsearch, once a field has been mapped as, say, text, you cannot later change it to keyword on an existing index — Lucene has already built the underlying segment files around the original type, and there is no in-place “change the column type” operation. The mapping for an existing field is frozen the moment the first document defines it (either explicitly, in your mapping definition, or implicitly, via dynamic mapping inferring a type from the first document Elasticsearch sees).

You can still evolve a mapping in two limited ways: adding entirely new fields to an existing mapping is always allowed, and some settings (like ignore_above on a keyword field) can be updated without touching existing data. But changing a field’s type, its analyzer, or converting a single value to an array in a way that breaks existing assumptions all require a full reindex. This is not a workaround or an edge case — reindexing is the supported migration path in Elasticsearch, and the tooling is built around it:

// 1. Create the new index with the corrected mapping
await es.indices.create({
  index: 'products_v2',
  body: {
    mappings: {
      properties: {
        // Example: correcting a mistake where `sku` was mapped as `text`
        // (tokenized, can't be used for exact lookups or aggregations)
        // and needs to become `keyword`.
        sku: { type: 'keyword' },
        name: { type: 'text', analyzer: 'product_search' },
        // ...rest of the mapping, corrected
      },
    },
  },
});

// 2. Reindex documents from the old index into the new one
await es.reindex({
  body: {
    source: { index: 'products' },
    dest: { index: 'products_v2' },
  },
});

// 3. Point an alias at the new index, and swap it atomically
await es.indices.updateAliases({
  body: {
    actions: [
      { remove: { index: 'products', alias: 'products_current' } },
      { add: { index: 'products_v2', alias: 'products_current' } },
    ],
  },
});

// 4. Once you've confirmed traffic is healthy against the alias,
//    delete the old index.
await es.indices.delete({ index: 'products' });

The alias indirection in step 3 is the detail that makes this safe in production: your application should never query a concrete index name directly (products) — it should query an alias (products_current) that you can repoint atomically once the reindex finishes and you’ve spot-checked the new data. If you bake the literal index name into your application code, every mapping fix becomes a coordinated deploy instead of a background job. This is worth setting up before you need your first reindex, not after — retrofitting alias indirection onto an application that already queries hard-coded index names is its own migration.

The practical takeaway for schema design: spend real time getting the mapping right before you load production data, because “we’ll just add a field later” is fine, but “we’ll just change this field’s type later” means a reindex of your entire dataset, which for a large index can take hours and requires double the disk space during the transition (old and new index coexisting).


// Index (create) a document
await es.index({
  index: 'products',
  id: '1',                    // Optional — auto-generated if omitted
  document: {
    name: 'Sony WH-1000XM5 Wireless Headphones',
    description: 'Industry-leading noise canceling with outstanding call quality',
    price: 349.99,
    category: 'Electronics',
    brand: 'Sony',
    rating: 4.8,
    inStock: true,
    tags: ['wireless', 'noise-canceling', 'headphones'],
    createdAt: new Date().toISOString(),
  },
});

// Bulk index — much faster for large imports
const operations = products.flatMap(product => [
  { index: { _index: 'products', _id: product.id } },
  product,
]);

await es.bulk({ operations });

// Get by ID
const doc = await es.get({ index: 'products', id: '1' });
console.log(doc._source);  // The document

// Update (partial)
await es.update({
  index: 'products',
  id: '1',
  doc: { price: 299.99, inStock: false },
});

// Delete
await es.delete({ index: 'products', id: '1' });

A detail worth internalizing here: Lucene segments are immutable once written, so the update call above is not a true in-place partial update at the storage layer — under the hood, Elasticsearch fetches the current _source, merges your changes into it, marks the old document as deleted, and indexes a brand-new document with the merged content. This is invisible from the client API, but it explains why high-frequency partial updates on the same document (a view counter incremented on every page load, for instance) generate far more I/O and segment churn than the “partial update” name suggests, and why a dedicated counter store (Redis, or a database with real row-level updates) is usually the better fit for that access pattern. It also explains why bulk indexing is not just a convenience API — batching writes into a single _bulk request amortizes the network round-trip and lets Elasticsearch’s indexing pipeline buffer more efficiently, which for large imports is easily an order of magnitude faster than issuing one index call per document.


Search Queries

// Simple match query
const results = await es.search({
  index: 'products',
  query: {
    match: {
      name: 'wireless headphones',   // Tokenized: ["wireless", "headphones"]
    },
  },
});

// Multi-field search with boosting
const results = await es.search({
  index: 'products',
  query: {
    multi_match: {
      query: 'wireless headphones',
      fields: ['name^3', 'description^1', 'tags^2'],  // name is 3x more important
      type: 'best_fields',
      fuzziness: 'AUTO',   // Handle typos: "headphoens" → "headphones"
    },
  },
});

Relevance ranking is where Elasticsearch stops resembling SQL in a way that trips up a lot of developers moving over from relational databases. In SQL, a WHERE clause is a boolean predicate: a row either satisfies it or it doesn’t, and there’s no notion of a row satisfying it “more” than another row. A match query, by contrast, returns every document that shares at least one token with the query and then ranks them by BM25 — a scoring function built from term frequency (how often the term appears in this document), inverse document frequency (how rare the term is across the whole index — matching on “the” scores far lower than matching on “headphones” because “the” appears everywhere and carries no discriminating signal), and field-length normalization (a match in a short field counts for more than the same match buried in a long one). The practical consequence is that “relevance tuning” in Elasticsearch is an ongoing exercise in adjusting field boosts, function scores, and analyzers to get results that feel right to end users — it is closer to product work than to writing a correct SQL predicate, and there generally isn’t a single “correct” ranking, only one that tests better for your users’ actual queries.

Bool Query — Combine Conditions

const results = await es.search({
  index: 'products',
  query: {
    bool: {
      must: [                          // All must match (AND)
        { match: { name: 'headphones' } },
      ],
      filter: [                        // Must match, no scoring impact
        { term: { inStock: true } },
        { term: { category: 'Electronics' } },
        {
          range: {
            price: { gte: 100, lte: 500 },
          },
        },
      ],
      should: [                        // Nice to have (boost score)
        { term: { brand: 'Sony' } },
        { range: { rating: { gte: 4.5 } } },
      ],
      must_not: [                      // Must NOT match
        { term: { brand: 'Unknown' } },
      ],
    },
  },
});

Notice that filter and must look similar but behave very differently, and mixing them up is a common source of both correctness bugs and slow queries. Clauses in filter participate in matching but not in scoring, and — critically — their results are cacheable by Elasticsearch across requests, because a filter like inStock: true either matches or it doesn’t regardless of the query text. Clauses in must affect the relevance score and generally cannot be cached the same way. If you put a condition that’s really a hard requirement (in stock, category = Electronics, price under a threshold) into must instead of filter, you don’t get a wrong answer, but you do throw away a real performance optimization and you pollute the relevance score with a condition that shouldn’t be affecting ranking at all.

The Analyzer Mismatch Trap

This is, in my experience, the single most confusing failure mode Elasticsearch produces, because the symptom looks like a bug in Elasticsearch itself: a search that returns results perfectly when you run it directly in Kibana’s Dev Tools returns nothing when the exact same query text comes from your application. Nothing in the error output points at the real cause, because there is no error — the query executes successfully and just doesn’t match.

The root cause is almost always that the field was indexed with one analyzer and the field is being queried with a different one — or the query is hitting a keyword sub-field when it should be hitting the analyzed text field, or vice versa. I ran into this on a project where the product_search custom analyzer (the one defined earlier in this guide, with lowercase, asciifolding, and a stop filter) was applied when the index was created, but the field mapping was later touched by a well-meaning teammate who added a keyword-style exact-match field with the same name as a multi-field without realizing the top-level field’s analyzer had silently reverted to the standard analyzer during a mapping merge. Kibana’s Dev Tools console was querying against a saved search that explicitly targeted the .keyword sub-field, so it worked. Our Node.js service was using multi_match against the base field, which was now tokenized differently than the documents had been indexed with — the stop-word filter meant “the wireless mouse” and “wireless mouse” produced different token streams than expected, and queries that should have matched simply didn’t.

The fix, once found, is a one-line mapping change, but finding it took longer than it should have because I was checking the query first, then the data, before finally checking the analyzers on both sides. The tool that actually resolves this quickly is _analyze, and it’s worth running proactively any time full-text search “isn’t working” rather than assuming the query DSL is wrong:

// Ask Elasticsearch: what tokens does this text produce when it
// goes through the analyzer configured for `name` on `products`?
const indexTimeTokens = await es.indices.analyze({
  index: 'products',
  field: 'name',
  text: 'Wireless Headphones',
});

// Compare against what a raw standard analyzer produces —
// if these two token lists differ, that's your mismatch.
const queryTimeTokens = await es.indices.analyze({
  analyzer: 'standard',
  text: 'Wireless Headphones',
});

console.log(indexTimeTokens.tokens.map(t => t.token));
console.log(queryTimeTokens.tokens.map(t => t.token));

The rule that comes out of this: whenever full-text search silently returns fewer results than expected (not zero, necessarily — sometimes just fewer, which is even harder to notice), check _analyze against the actual field before touching the query. It takes thirty seconds and rules out an entire category of bugs that otherwise costs an afternoon.

Pagination and Sorting

const results = await es.search({
  index: 'products',
  query: { match: { name: 'headphones' } },
  from: 0,           // Offset (page - 1) * size
  size: 10,          // Page size
  sort: [
    { _score: 'desc' },         // Primary: relevance score
    { rating: 'desc' },         // Secondary: highest rated first
    { price: 'asc' },           // Tertiary: cheapest first
  ],
  _source: ['name', 'price', 'rating', 'category'],  // Only these fields
});

const hits = results.hits.hits;
const total = results.hits.total.value;

Aggregations

Aggregations compute summaries and analytics:

const results = await es.search({
  index: 'products',
  query: {
    bool: {
      filter: [{ term: { inStock: true } }],
    },
  },
  size: 0,                 // Only return aggregations, no documents
  aggs: {
    // Category breakdown
    by_category: {
      terms: { field: 'category', size: 10 },
      aggs: {
        avg_price: { avg: { field: 'price' } },
        avg_rating: { avg: { field: 'rating' } },
      },
    },

    // Price histogram
    price_ranges: {
      histogram: {
        field: 'price',
        interval: 100,     // Buckets: 0-100, 100-200, 200-300...
      },
    },

    // Statistics
    price_stats: {
      stats: { field: 'price' },  // min, max, avg, count, sum
    },

    // Top brands
    top_brands: {
      terms: { field: 'brand', size: 5 },
    },
  },
});

// Access results
const categories = results.aggregations.by_category.buckets;
// [{ key: 'Electronics', doc_count: 245, avg_price: { value: 187.5 } }, ...]

Aggregations are the feature that most clearly separates Elasticsearch from a plain search engine and puts it into “analytics engine” territory — a single request above computes a faceted category breakdown, a price histogram, and summary statistics, all over the same filtered document set, in one round trip. The equivalent in SQL would be several separate GROUP BY queries (or a more elaborate GROUPING SETS construct), each re-scanning the filtered rows. Because Elasticsearch pre-computes term statistics as part of its indexing pipeline, these aggregations run against the inverted index structure rather than a full document scan, which is part of why faceted search UIs (the category/price/brand sidebar filters you see on e-commerce sites) are a natural fit for Elasticsearch and painful to build efficiently on a relational database at scale.


Shard Sizing: The Trade-off Nobody Tells You About

Sharding is where Elasticsearch’s distributed-by-default design becomes a source of real operational pain if you get it wrong early, and unlike the mapping mistakes above, a bad shard count is also something you can only fix by reindexing — number_of_shards is set at index creation time and cannot be changed afterward (number_of_replicas can be changed anytime; shard count cannot).

The trade-off runs in both directions, and neither extreme is free:

  • Too many small shards. Each shard is a separate Lucene index with its own file handles, memory overhead (segment metadata, field data caches), and per-shard query overhead — Elasticsearch has to fan a query out to every shard and merge the results, even if a shard holds a handful of documents. A cluster with thousands of tiny shards spends a disproportionate amount of its resources on coordination overhead rather than actual search work, and can hit low-level limits (open file descriptors, JVM heap pressure from cluster state) well before it runs out of disk.
  • Too few, oversized shards. A single shard is the unit of relocation and recovery — if a node holding a 200GB shard fails, recovering that shard (either from replicas or from a snapshot) takes proportionally longer than recovering several smaller shards in parallel, and a shard that’s too large to fit comfortably in a node’s available heap and OS page cache will make queries against it noticeably slower. Very large shards also make rebalancing across the cluster slow and disruptive, because a shard can only move as a whole unit — you can’t split a load across two nodes for one 300GB shard.

I’ve watched this go wrong in the too-many-small-shards direction specifically: a project created a new daily index per tenant per day (a fairly common time-series pattern), each with the default shard count from an older Elasticsearch version, on a dataset where most tenants had only a few hundred documents a day. Within a few months the cluster had accumulated many thousands of shards, the vast majority holding a trivial number of documents, and the cluster’s health checks started timing out — not because any single query was slow, but because cluster state updates (which every shard change has to propagate) and query fan-out overhead had scaled with shard count, not with data volume. The fix was consolidating to weekly or monthly indices with a shard count sized to actual data volume, plus setting index.number_of_shards explicitly per index template instead of relying on defaults — but that meant a full reindex of historical data, which is exactly the kind of expensive migration you want to avoid by getting the sizing right up front.

A reasonable starting heuristic: aim for shards in the tens-of-GB range (commonly cited guidance is roughly 10–50GB per shard, though this varies with your query patterns and hardware), size number_of_shards based on your projected data volume rather than what you have on day one, and prefer time-based index patterns (daily/weekly/monthly) with an index lifecycle management policy over one enormous index, since it lets you age out and delete old data by dropping whole indices instead of running expensive delete-by-query operations.


Refresh Interval: Near-Real-Time vs Indexing Throughput

One more setting that’s easy to overlook until it causes a support ticket: by default, Elasticsearch refreshes each shard every second (index.refresh_interval: "1s"), which is what people mean when they call Elasticsearch “near-real-time” rather than truly real-time — a document you just indexed is not searchable until the next refresh happens, not the instant es.index() returns. A refresh is not free: it forces Lucene to write a new segment and open a new searcher, and if you’re doing a large bulk load, refreshing every second means Elasticsearch is constantly creating small segments that then need to be merged together later, which competes with your bulk indexing for I/O and CPU.

The practical pattern for a bulk import is to raise the refresh interval (or disable it entirely) for the duration of the load, then restore it once you need the data queryable in near-real-time again:

// Before a large bulk import: reduce refresh overhead
await es.indices.putSettings({
  index: 'products',
  body: { index: { refresh_interval: '-1' } },  // disable periodic refresh
});

await es.bulk({ operations: largeBatchOfOperations });

// Restore near-real-time search once the bulk load is done
await es.indices.putSettings({
  index: 'products',
  body: { index: { refresh_interval: '1s' } },
});

Conversely, if your use case genuinely doesn’t need sub-second searchability — a nightly analytics index that’s queried the next morning, for instance — deliberately setting a longer refresh interval (30s, or even disabling it and calling _refresh manually after a load) reduces steady-state overhead with no user-visible cost. The mistake to avoid is treating refresh_interval as a knob you set once and forget: it should match the actual latency requirement of the use case, and for most transactional-feeling search UIs, the 1-second default is the right trade-off between freshness and indexing throughput.


Autocomplete / Suggestions

// Using completion suggester (defined in mapping as 'completion' type)
const suggestions = await es.search({
  index: 'products',
  suggest: {
    product_suggest: {
      prefix: 'wirel',            // What the user has typed so far
      completion: {
        field: 'name.suggest',    // The completion field
        size: 5,
        skip_duplicates: true,
        fuzzy: { fuzziness: 1 }, // Allow 1 character difference
      },
    },
  },
});

const options = suggestions.suggest.product_suggest[0].options;
// [{ text: 'Wireless Headphones', _score: 1 }, ...]

Node.js Integration Pattern

// services/search.ts
import { es } from '../lib/elasticsearch';

interface SearchParams {
  query: string;
  category?: string;
  minPrice?: number;
  maxPrice?: number;
  inStock?: boolean;
  page?: number;
  pageSize?: number;
  sortBy?: 'relevance' | 'price_asc' | 'price_desc' | 'rating';
}

export async function searchProducts(params: SearchParams) {
  const {
    query,
    category,
    minPrice,
    maxPrice,
    inStock,
    page = 1,
    pageSize = 20,
    sortBy = 'relevance',
  } = params;

  const sort = {
    relevance: [{ _score: 'desc' }],
    price_asc: [{ price: 'asc' }],
    price_desc: [{ price: 'desc' }],
    rating: [{ rating: 'desc' }],
  }[sortBy];

  const filters: any[] = [];
  if (category) filters.push({ term: { category } });
  if (inStock !== undefined) filters.push({ term: { inStock } });
  if (minPrice !== undefined || maxPrice !== undefined) {
    filters.push({
      range: {
        price: {
          ...(minPrice !== undefined && { gte: minPrice }),
          ...(maxPrice !== undefined && { lte: maxPrice }),
        },
      },
    });
  }

  const results = await es.search({
    index: 'products',
    query: {
      bool: {
        must: query
          ? [{
              multi_match: {
                query,
                fields: ['name^3', 'description^1', 'tags^2'],
                fuzziness: 'AUTO',
              },
            }]
          : [{ match_all: {} }],
        filter: filters,
      },
    },
    sort,
    from: (page - 1) * pageSize,
    size: pageSize,
    aggs: {
      by_category: { terms: { field: 'category', size: 20 } },
      price_stats: { stats: { field: 'price' } },
    },
  });

  return {
    products: results.hits.hits.map(hit => ({ id: hit._id, ...hit._source })),
    total: results.hits.total.value,
    totalPages: Math.ceil(results.hits.total.value / pageSize),
    facets: {
      categories: results.aggregations.by_category.buckets,
      priceStats: results.aggregations.price_stats,
    },
  };
}

Syncing with a Primary Database

// Keep Elasticsearch in sync when the primary DB changes
import { prisma } from '../lib/prisma';
import { es } from '../lib/elasticsearch';

async function indexProduct(productId: number) {
  const product = await prisma.product.findUnique({
    where: { id: productId },
    include: { category: true, tags: true },
  });

  if (!product) return;

  await es.index({
    index: 'products',
    id: product.id.toString(),
    document: {
      name: product.name,
      description: product.description,
      price: product.price,
      category: product.category.name,
      brand: product.brand,
      rating: product.rating,
      inStock: product.stock > 0,
      tags: product.tags.map(t => t.name),
      createdAt: product.createdAt.toISOString(),
    },
  });
}

async function deleteProductIndex(productId: number) {
  await es.delete({ index: 'products', id: productId.toString() });
}

// Call after DB operations
await prisma.product.update({ where: { id }, data: updateData });
await indexProduct(id);   // Sync to Elasticsearch

This dual-write pattern — write to Postgres, then write to Elasticsearch — is the simplest way to keep the two in sync, and it’s genuinely fine for a lot of applications, but it’s worth being honest about its failure mode: if the process crashes, times out, or the Elasticsearch write fails after the database commit succeeds, the two stores drift out of sync silently. There’s no transaction spanning both systems, so you can’t roll one back to match the other. For low-stakes search indexes (a “recently updated” flag being briefly stale) this is an acceptable risk. For anything where staleness is actually costly — inventory-sensitive search results, for example — the more robust pattern is change-data-capture: a tool like Debezium tails the database’s write-ahead log and publishes change events to a queue, with a separate consumer responsible for updating Elasticsearch. That adds real infrastructure (Kafka or an equivalent), so it’s a trade-off worth making deliberately rather than defaulting into either direction — starting with dual writes and a periodic reconciliation job (a scheduled task that diffs a sample of documents between Postgres and Elasticsearch and re-indexes mismatches) is often the pragmatic middle ground before a full CDC pipeline is justified.


Quick Reference

OperationCode
Create indexes.indices.create({ index, body: { settings, mappings } })
Index documentes.index({ index, id, document })
Bulk indexes.bulk({ operations })
Full-text search{ match: { field: 'query' } }
Exact match{ term: { field: 'value' } }
Range filter{ range: { price: { gte: 10, lte: 100 } } }
Multiple conditions{ bool: { must, filter, should, must_not } }
Aggregate{ aggs: { name: { terms: { field: 'category' } } } }

Frequently Asked Questions (FAQ)

Q. Why doesn’t a document I just indexed show up when I search for it right away?

A. Elasticsearch is near-real-time: a newly indexed document becomes searchable only after the next shard refresh, which runs every second by default (index.refresh_interval: "1s") and not at all if you set it to -1 for a bulk load. This often shows up in integration tests that index and then immediately search. For those cases, pass refresh: 'wait_for' on the index request or call _refresh explicitly, but avoid forcing a refresh on every write in production, because it creates many small segments that compete with indexing for I/O.