Files
summercms/docs/services/search.md
Jakub Zych 3ae49bb6f4 feat(12-01): report the search engine's found count and field weights
- optional beachcomber.PageSearcher returns a page of candidate ids plus
  the engine's found count; beachcomber.SearchPage falls back to
  SearchIDs for engines without it, so Engine is unchanged
- Query.QueryByWeights is sent to Typesense as query_by_weights; a
  mismatched weight list or a page above typesense.MaxPerPage (250) is
  refused before any request
2026-10-02 11:37:01 +02:00

9.5 KiB

title, description, section, order
title description section order
Search Keep models in a search index with beachcomber, synced after commit behind a kill-switch, and re-check the candidate IDs a search returns in SQL. services 140

Search

beachcomber is the SummerCMS counterpart of Laravel Scout, as WinterCMS applications use it without a queue. A model opts in by implementing beachcomber.Searchable; after a write of such a model commits, its row is reloaded and its document written to, or removed from, the search index. A search asks the index for matching IDs, and the application loads the rows from the database.

Engines

search.driver selects the engine. The default, null, indexes nothing and finds nothing. The typesense engine, from the beachcomber/typesense package, talks to a Typesense server over its HTTP API. An engine registers itself from its package's init function, so the application imports the engine package for its side effect; an unknown name stops the start-up. search.prefix is prepended to every index name, for example to keep staging and production apart on one server:

cfg, err := compass.Open(compass.Options{Dir: "config", Env: "development", Environ: []string{}})
if err != nil {
	fmt.Println(err)
	return
}
_ = cfg.Set("search.prefix", "staging_")
svc, err := beachcomber.From(backpack.New(cfg)) // search.driver defaults to null
if err != nil {
	fmt.Println(err)
	return
}
fmt.Println(svc.Engine().Name(), svc.Engine().Configured(), svc.IndexName(&Post{}))
ids, err := svc.Engine().SearchIDs(context.Background(), svc.IndexName(&Post{}), beachcomber.Query{Q: "go"})
fmt.Println(ids, err)

_ = cfg.Set("search.driver", "elastic")
_, err = beachcomber.From(backpack.New(cfg))
fmt.Println(err != nil)
// Output:
// null false staging_acme_blog_posts
// [] <nil>
// true

An engine implements beachcomber.Engine and registers with beachcomber.RegisterEngine. The examples on this page use a small in-memory engine that matches titles:

func (e *memoryEngine) SearchIDs(ctx context.Context, index string, q beachcomber.Query) ([]string, error) {
	e.mu.Lock()
	defer e.mu.Unlock()
	ids := []string{}
	for id, d := range e.docs[index] {
		if strings.Contains(strings.ToLower(fmt.Sprint(d["title"])), strings.ToLower(q.Q)) {
			ids = append(ids, id)
		}
	}
	slices.Sort(ids)
	return ids, nil
}

Searchable models

A model implements beachcomber.Searchable with three methods, and needs no import of beachcomber to do so:

// SearchableAs is the index name, before search.prefix.
func (Post) SearchableAs() string { return "acme_blog_posts" }
// ShouldBeSearchable keeps drafts out of the index.
func (p *Post) ShouldBeSearchable() bool { return p.Published }
// ToSearchableArray builds the document from the committed row.
func (p *Post) ToSearchableArray(ctx context.Context, db *gorm.DB) (map[string]any, error) {
	return map[string]any{
		"id":      strconv.FormatUint(uint64(p.ID), 10),
		"blog_id": int64(p.BlogID),
		"title":   p.Title,
	}, nil
}

ToSearchableArray runs after the commit on a fresh copy of the row, so it may query related rows. A row whose ShouldBeSearchable is false, a soft-deleted row and a deleted row have their documents removed. A model can also implement beachcomber.IndexSchemaProvider, the schema the engine creates a missing index with, and beachcomber.SearchKeyer, a document key other than the primary key.

Sync after commit

beachcomber.From installs GORM callbacks that register the sync with lagoon.AfterCommit:

  • Inside lagoon.Transaction, the sync runs after the commit, and never after a rollback.
  • A single-statement write syncs after GORM commits it.
  • Inside a plain GORM transaction the sync is skipped with a warning, because the commit cannot be observed. Wrap such writes in lagoon.Transaction, or call beachcomber.Service.Sync after the commit.

The sync runs inline in the writing goroutine, so a search right after a save finds the document, and it is bounded by the engine's timeout. It is never fatal: a failure is logged as search: sync failed with the index, key and operation, never the document or the API key, and the write stays committed. A write without a primary key value, such as Model(&Post{}).Where(...).Updates(...), cannot be synced row by row; bulk paths call beachcomber.Service.Sync and beachcomber.Service.Remove per row, or reindex.

The kill-switch

Nothing is sent when the engine is not configured (the null engine, or Typesense without an API key), when no database is published, or when the application's beachcomber.Gate reports off. Install the gate from a plugin's Boot. A gate must treat a read error as off:

svc.SetGate(beachcomber.GateFunc(func(ctx context.Context, db *gorm.DB) bool {
	var enabled bool
	err := db.WithContext(ctx).Raw(`SELECT search_enabled FROM acme_blog_settings WHERE id = 1`).Scan(&enabled).Error
	return err == nil && enabled
}))

Searching

beachcomber.Engine.SearchIDs returns the IDs of matching documents, in the engine's order. They are candidates, not answers: the index can be stale (a write it missed, a document from before a permission change) and its filters are only as good as the document. Re-check every ID in SQL, with the same ownership, visibility and soft-delete conditions the rest of the API applies, before a row reaches a response:

ids, err := svc.Engine().SearchIDs(ctx, svc.IndexName(&Post{}), beachcomber.Query{
	Q:        term,
	QueryBy:  []string{"title"},
	FilterBy: "blog_id:=" + strconv.FormatUint(uint64(blogID), 10),
})
if err != nil {
	return nil, err
}
// The ids are candidates from an index that may be stale or loosely
// filtered: re-check every one in SQL before exposing a row.
var posts []Post
err = db.WithContext(ctx).
	Where("id IN ? AND blog_id = ? AND published", ids, blogID).
	Order("id").
	Find(&posts).Error
return posts, err

An empty result is an empty list, never an error.

Warning

Never return rows, or even counts, straight from search IDs. A stale index would otherwise show a draft, a deleted record or another user's data.

A paginated endpoint also needs to know how many documents matched. An engine that can tell implements the optional beachcomber.PageSearcher interface; call it through beachcomber.SearchPage, which returns a beachcomber.SearchResult with the page of IDs and Found, the number of documents the engine matched. For an engine without PageSearcher the helper calls SearchIDs and sets Found to the number of IDs returned, so code written against it works with every engine. Found is a candidate count like the IDs: it includes stale and foreign documents. Use it to decide how many IDs to fetch, then recount the total in SQL with the same conditions as the rows, as Laravel Scout does when a query callback is set.

beachcomber.Query.QueryByWeights ranks the fields of QueryBy, one weight per field in the same order, as Scout's query_by_weights option does.

Typesense

The Typesense engine follows the Scout Typesense wire contract, so indexes built by a WinterCMS application can be searched by the port:

driver: typesense
typesense:
  host: 127.0.0.1
  port: 8108
  protocol: http

in config/search.yaml, with the key in SUMMER_SEARCH__TYPESENSE__API_KEY. Without an API key nothing is ever sent. A search is one request to the collection's search endpoint, and the engine returns the hit IDs in Typesense's order:

// A stand-in Typesense node.
node := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
	q := r.URL.Query()
	fmt.Println(r.Method, r.URL.Path, "key sent:", r.Header.Get("X-TYPESENSE-API-KEY") != "")
	fmt.Println("q", q.Get("q"), "query_by", q.Get("query_by"), "filter_by", q.Get("filter_by"))
	fmt.Fprint(w, `{"hits":[{"document":{"id":"3"}},{"document":{"id":"1"}}]}`)
}))
defer node.Close()
u, _ := url.Parse(node.URL)
port, _ := strconv.Atoi(u.Port())

engine := typesense.New(typesense.Config{
	APIKey:            "test-only-key", // SUMMER_SEARCH__TYPESENSE__API_KEY
	Host:              u.Hostname(),
	Port:              port,
	Protocol:          "http",
	ConnectionTimeout: 2 * time.Second,
})
ids, err := engine.SearchIDs(context.Background(), "acme_blog_posts", beachcomber.Query{
	Q:        "go",
	QueryBy:  []string{"title"},
	FilterBy: "blog_id:=7",
})
fmt.Println(ids, err)

// Without an API key the engine is not configured and nothing is sent.
fmt.Println(typesense.New(typesense.Config{}).Configured())
// Output:
// GET /collections/acme_blog_posts/documents/search key sent: true
// q go query_by title filter_by blog_id:=7
// [3 1] <nil>
// false

The Typesense engine implements beachcomber.PageSearcher on the same request, adding the answer's found count, and sends query_by_weights when the query has weights. A page is at most typesense.MaxPerPage (250) documents, Typesense's own limit: a larger PerPage is an error, so fetch more IDs page by page. Each request times out after search.typesense.connection_timeout_seconds (2 seconds by default). A failed answer is a typesense.StatusError with the method, path and status, never the answer body.