mirror of
https://github.com/basicmachines-co/basic-memory
synced 2026-06-21 13:47:35 +00:00
fix(core): Postgres parity for FTS punctuation/relaxation (CI green)
CI Postgres shard caught two issues invisible to the local SQLite suite:
1. Postgres _prepare_single_term regression: the new edge-punctuation
strip ran after special-character cleaning, so an all-special-char
term ("()&!:") collapsed to empty and skipped the existing
NOSPECIALCHARS:* guard, emitting a malformed ":*". Folded the strip
into the word handlers so every guard survives, and added a
single-word empty guard.
2. Backend-specific test assumptions. Four tests in
test_search_repository.py (run under both backends via the
search_repository fixture) asserted SQLite FTS5 syntax and
SQLite-only strict-miss behavior. Postgres to_tsquery('english', ...)
auto-strips stopwords, so "When did Melanie paint a sunrise?" already
matches under strict AND. Made the four tests backend-aware via the
existing is_postgres_backend() helper, and switched the relaxation
integration test to a query with a word absent from the doc
("hiking") so the strict miss holds on both backends.
Reproduced and fixed against real Postgres (testcontainers): full
search test surface green on both backends (53 passed Postgres,
2968 SQLite), ruff + ty clean.
Signed-off-by: Drew Cain <groksrc@gmail.com>
This commit is contained in:
@@ -243,17 +243,14 @@ class PostgresSearchRepository(SearchRepositoryBase):
|
||||
for char in special_chars:
|
||||
cleaned_term = cleaned_term.replace(char, " ")
|
||||
|
||||
# Sentence punctuation carries no lexical signal in tsquery either;
|
||||
# strip it from word edges so question-form queries produce clean
|
||||
# lexemes (parity with the SQLite FTS5 preparation).
|
||||
if " " in cleaned_term:
|
||||
cleaned_term = " ".join(
|
||||
word.strip("?!.,;") for word in cleaned_term.split() if word.strip("?!.,;")
|
||||
)
|
||||
|
||||
# Handle multi-word queries
|
||||
if " " in cleaned_term:
|
||||
words = [w for w in cleaned_term.split() if w.strip()]
|
||||
# Strip sentence punctuation from word edges so question-form
|
||||
# queries produce clean lexemes (parity with SQLite FTS5 prep).
|
||||
# The tsquery tokenizer ignores this punctuation anyway; leaving it
|
||||
# in only risks tsquery syntax errors. Interior characters are kept.
|
||||
words = [w.strip("?!.,;") for w in cleaned_term.split()]
|
||||
words = [w for w in words if w]
|
||||
if not words:
|
||||
# All characters were special chars, search won't match anything
|
||||
# Return a safe search term that won't cause syntax errors
|
||||
@@ -266,8 +263,11 @@ class PostgresSearchRepository(SearchRepositoryBase):
|
||||
# Join with AND operator
|
||||
return " & ".join(prepared_words)
|
||||
|
||||
# Single word
|
||||
cleaned_term = cleaned_term.strip()
|
||||
# Single word: strip edge punctuation; guard the now-empty case so a
|
||||
# bare ":*"/"" never reaches tsquery.
|
||||
cleaned_term = cleaned_term.strip().strip("?!.,;")
|
||||
if not cleaned_term:
|
||||
return "NOSPECIALCHARS:*"
|
||||
if is_prefix:
|
||||
return f"{cleaned_term}:*"
|
||||
else:
|
||||
|
||||
Reference in New Issue
Block a user