2026-10-11 17:10 UTC

SeaSearch's maintainers claim their newly open-sourced Go engine combines Elasticsearch-compatible full-text and vector search with S3-backed shared storage and per-tenant indexes, reducing the footprint and cluster-management burden of multi-tenant retrieval.

state: seedheat: lowuncertainty: mediumnovelscott: lowretrieval-systems knowledge-systems self-hostingDaniel PanSeafileseacloud-lab

What is this?

SeaSearch is presented in Reddit announcement snippets as a Go-based, S3-backed search engine for multi-tenant applications, with shared index storage intended to reduce infrastructure overhead. The supplied case describes a new open-source release associated with Daniel Pan, Seafile, and seacloud-lab, but the web snippets do not establish their roles or verify the release. Elasticsearch API compatibility, combined full-text/vector indexing, and reduced resource requirements remain case-reported claims; the snippets provide no compatibility tests or benchmarks.

Why it matters to Scott

Scott’s search project uses ChromaDB, while his appliance architecture isolates complete stacks per customer; the hits establish no Elasticsearch dependency or shared multi-tenant indexing burden that SeaSearch would address, and its claimed savings remain unverified. The radar already tracks related object-storage retrieval designs in Polign and Zep, but not this SeaSearch development; no substantive challenge to or independent adoption of Scott’s positions is established.
radar:polign-stateless-agent-memoryradar:zep-konig-graph-memory-scalingradar:concept.retrieval-infrastructure
queries asked of Scott's wikis
  • Multi-tenant retrieval per-tenant indexes and isolation
  • Hybrid full-text vector search in knowledge systems
  • Self-hosted search operational complexity and resource budgets
  • Object-storage-backed indexes storage compute separation
  • Elasticsearch compatibility retrieval backend migration

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 613h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-16 03:24 (minted)⭐ origin echo-reconstructedSeaSearch is presented as a lightweight Go search engine with Elasticsearch API compatibility, full-text and vector indexing, and shared S3
seacloud-lab on github (echo) · attributed from hn.story.49721546 · published time unknown
—
09-16 02:45first on hacker news · published · lag ?Show HN: SeaSearch – Lightweight, S3-backed multi-tenant search engine
Daniel-Pan
—
09-16 02:45amplified on hacker news 👑hn.story.49721546
Daniel-Pan
peak 2 · 2 comments · 98% of case engagement
09-16 03:20our radar first saw it · lag ?discovery anchor: hn.story.49721546—
pace: p36 vs 1032 stories at the 336h mark (now 613h old) — ahead of agentgate-signed-agent-receipts (1.3x), behind agent-memory-add-search-evaluation (0.8x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: SeaSearch – Lightweight, S3-backed multi-tenant search engine
Retrieved article excerpt

Open article · Retrieved 2026-09-16T03:21:48.249148+00:00

# SeaSearch — Lightweight, Go-based multi-tenant search engine

**SeaSearch** is a lightweight, Go-based multi-tenant search engine featuring Elasticsearch API compatibility and S3-backed storage—designed to support unlimited indexes without overhead.

In a multi-tenant environment, such as a SaaS application, this allows each tenant's data to be indexed independently. With traditional search engines such as Elasticsearch, when all tenants' data is stored in a single index, the index may eventually become too large and require manual sharding. With SeaSearch, each tenant can have its own index, making it easier to manage and scale large numbers of tenants.

## SeaSearch vs. Elasticsearch

- **Lightweight**: SeaSearch is implemented in Go and has a smaller runtime footprint than Elasticsearch, which is built on the JVM.
- **No Practical Limit on the Number of Indexes**: SeaSearch is designed to support a large number of indexes. This makes it possible to create a separate index for each tenant, project, or other logical unit in an application. Queries can then be restricted to the relevant index, reducing the amount of data that needs to be searched. With Elasticsearch, applications often store data from many tenants or projects in the same index, which can become less efficient as the data volume grows.
- **Elasticsearch API Compatibility**: SeaSearch provides an API compatible with Elasticsearch, making it easier to integrate with existing applications.
- **S3-Compatible Storage**: SeaSearch can use S3-compatible object storage as its storage backend.
- **Shared-Storage Cluster Architecture**: Elasticsearch clusters replicate data across nodes, which can make cluster management and scaling relatively complex. SeaSearch uses a shared-storage architecture in which cluster nodes share the same storage backend, typically S3-compatible object storage. This simplifies cluster management and makes it easier to provide high availability. Query performance can also be scaled horizontally by adding more query nodes.
- **Vector Search**: SeaSearch provides a lightweight vector search implementation with support for Flat, HNSW, and IVFPQ vector indexes.

## Architecture

SeaSearch uses a shared-storage architecture.

[image](https://private-user-images.githubusercontent.com/109284/647158240-cafead47-ec61-4c29-86f3-5617ee1f57b0.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3ODk1MjkyMDcsIm5iZiI6MTc4OTUyODkwNywicGF0aCI6Ii8xMDkyODQvNjQ3MTU4MjQwLWNhZmVhZDQ3LWVjNjEtNGMyOS04NmYzLTU2MTdlZTFmNTdiMC5wbmc_WC1BbXotQWxnb3JpdGhtPUFXUzQtSE1BQy1TSEEyNTYmWC1BbXotQ3JlZGVudGlhbD1BS0lBVkNPRFlMU0E1M1BRSzRaQSUyRjIwMjYwOTE2JTJGdXMtZWFzdC0xJTJGczMlMkZhd3M0X3JlcXVlc3QmWC1BbXotRGF0ZT0yMDI2MDkxNlQwMzIxNDdaJlgtQW16LUV4cGlyZXM9MzAwJlgtQW16LVNpZ25hdHVyZT0wMzEyOGY2NjkzYzhjMGRhODkwYmZmNDY0Zjc0NjMzMGFlN2VjZDYzZjJjYzI0NjU4OGViNzYwZDVmYzJkNWRlJlgtQW16LVNpZ25lZEhlYWRlcnM9aG9zdCZyZXNwb25zZS1jb250ZW50LXR5cGU9aW1hZ2UlMkZwbmcifQ.DnhS_wJUO9gNGEWuUtNxL_1vqooPI7GPZo4AKMPCqK8)

A SeaSearch cluster consists of the following types of nodes:

- **SeaSearch compute node**: Handles index read and write requests. All compute nodes share the same S3-compatible storage backend, where the index data is stored.
- **SeaSearch proxy (or gateway)**: Distributes client requests among SeaSearch compute nodes.
- **Etcd**: Stores cluster metadata, including index metadata and the data distribution map.
- **SeaSearch cluster manager**: Monitors the health of SeaSearch compute nodes and redistributes index ownership among the nodes when necessary.

In a single-node deployment, SeaSearch uses a local KV database (bbolt) to store index metadata and the local file system to store index data.

### Data Distribution and Failover

Because all compute nodes share the same storage backend, SeaSearch only needs to distribute **index ownership** among nodes rather than moving or replicating the actual index data.

- All indexes are grouped into a fixed number of partitions based on a hash of their names.
- The cluster manager maintains a map that determines which node currently owns each index partition. This map is stored in Etcd.
- The SeaSearch proxy routes requests for an index to the compute node that owns the corresponding partition.
- Whenever a compute node fails or a new node is added, the cluster manager recomputes the ownership map and transfers partition ownership between nodes as needed.

Because updating the cluster configuration does not require transferring large amounts of data, SeaSearch can efficiently manage a large number of indexes.

### Local Cache

When handling requests, compute nodes may need to retrieve index data from S3 storage, which can introduce additional latency. To reduce this latency, SeaSearch compute nodes cache index data on their local disks.

This caching strategy is feasible because index data is organized into immutable segments. Once created, an index segment can only be read or deleted; its contents are never modified.

To support queries against indexes that are larger than the available local disk space, compute nodes use a rotating cache. When a new segment needs to be cached and the cache has reached its size limit, older segments are evicted to make room.

With this design, clients typically experience higher latency only for the first request after an index segment has been evicted or when a node starts up. Once the cache is warmed up, subsequent requests can be served at speeds comparable to those of local storage.

In our experience, the warm-up latency can be further reduced by taking advantage of the high network bandwidth available in modern data centers. During the warm-up stage, multiple index segments can be retrieved from S3 in parallel, significantly accelerating index loading.

**Distributed Query Execution**: To further improve the ability to serve queries against very large indexes, SeaSearch can automatically distribute a search query across multiple compute nodes. Each node loads and searches a portion of the index data in parallel, and the results are then aggregated. This approach not only accelerates query execution but also reduces cache pressure on individual nodes, allowing SeaSearch to efficiently serve indexes that are significantly larger than the local disk capacity of a single node.
Daniel-Pan22
🟧 echo.github ⭐SeaSearch is presented as a lightweight Go search engine with Elasticsearch API compatibility, full-text and vector indexing, and shared S3 seacloud-lab——

Interpretation history

Decision trace