mirror of
https://github.com/basicmachines-co/basic-memory
synced 2026-06-21 13:47:35 +00:00
Compare commits
5 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| b5f13d6903 | |||
| 6d06c4a4e7 | |||
| 16869867da | |||
| 18da6194f9 | |||
| a6d0784335 |
@@ -116,22 +116,24 @@ After PyPI release is published, update the MCP registry:
|
||||
|
||||
#### Website Updates
|
||||
|
||||
**1. basicmachines.co** (sibling `basicmachines.co` repo)
|
||||
- **Goal**: Update version number displayed on the homepage
|
||||
- **Location**: Search for "Basic Memory v0." in the codebase to find version displays
|
||||
- **What to update**:
|
||||
- Hero section heading that shows "Basic Memory v{VERSION}"
|
||||
- "What's New in v{VERSION}" section heading
|
||||
- Feature highlights array (look for array of features with title/description)
|
||||
- **Process**:
|
||||
**1. basicmemory.com** (sibling `basicmemory.com` repo —
|
||||
`basicmachines-co/basicmemory.com`, formerly `basicmachines.co`)
|
||||
- **No version bump needed.** The marketing site is an Astro + React app and
|
||||
carries **no hardcoded Basic Memory version number** anywhere in its UI
|
||||
(`hero.tsx` and the rest of the site have no version string). The old
|
||||
instruction to bump `src/components/sections/hero.tsx` is obsolete — that
|
||||
file no longer holds a version. Release announcements are dated blog posts,
|
||||
not an in-place edit.
|
||||
- **Skip entirely for patch releases.**
|
||||
- **Significant releases only — optional announcement post**:
|
||||
1. Pull latest from GitHub: `git pull origin main`
|
||||
2. Create release branch: `git checkout -b release/v{VERSION}`
|
||||
3. Search codebase for current version number (e.g., "v0.16.1")
|
||||
4. Update version numbers to new release version
|
||||
5. Update feature highlights with 3-5 key features from this release (extract from CHANGELOG.md)
|
||||
6. Commit changes: `git commit -m "chore: update to v{VERSION}"`
|
||||
7. Push branch: `git push origin release/v{VERSION}`
|
||||
- **Deploy**: Follow deployment process for basicmachines.co
|
||||
3. Add a dated post under `src/content/blog/` modeled on an existing
|
||||
release post (e.g. `basic-memory-v0-19-0-release.md`), summarizing 3–5
|
||||
headline features from `CHANGELOG.md`
|
||||
4. Commit (`git commit -s -m "..."`), push, and open a PR against
|
||||
`basicmachines-co/basicmemory.com`
|
||||
- **Deploy**: follow that repo's deployment process.
|
||||
|
||||
**2. docs.basicmemory.com** (sibling `docs.basicmemory.com` repo)
|
||||
- **Goal**: Add a What's New page for the release and bump the homepage badge
|
||||
|
||||
@@ -322,7 +322,12 @@ The recipe runs `just lint` + `just typecheck`, then updates every release manif
|
||||
|
||||
**Post-release tasks** the recipe surfaces but doesn't run:
|
||||
- `docs.basicmemory.com` — add a What's New page under `content/2.whats-new/` and bump the version badge in `content/index.md` (the changelog page auto-fetches GitHub releases; see that repo's CLAUDE.md version-bump checklist)
|
||||
- `basicmachines.co` — bump version in `src/components/sections/hero.tsx`
|
||||
- `basicmemory.com` — the marketing site (Astro + React, repo
|
||||
`basicmachines-co/basicmemory.com`, formerly `basicmachines.co`) carries **no
|
||||
hardcoded version number** in its UI, so there is nothing to bump. For a
|
||||
significant release, optionally add a dated announcement post under
|
||||
`src/content/blog/` (model it on an existing `basic-memory-vX-Y-Z-release.md`).
|
||||
Skip entirely for routine patch releases.
|
||||
- MCP Registry — `mcp-publisher publish` from the repo root
|
||||
|
||||
See `.claude/commands/release/release.md` (and `beta.md`, `release-check.md`, `changelog.md` alongside it) for the full release + post-release runbook, including the slash commands.
|
||||
|
||||
@@ -6,6 +6,7 @@
|
||||
[](https://github.com/astral-sh/ruff)
|
||||

|
||||

|
||||
[](https://deepwiki.com/basicmachines-co/basic-memory)
|
||||
|
||||
## Skip the install — try Basic Memory in the cloud
|
||||
|
||||
|
||||
@@ -468,7 +468,8 @@ release version:
|
||||
echo "📝 REMINDER: Post-release tasks:"
|
||||
echo " 1. docs.basicmemory.com - Add a What's New page under content/2.whats-new/"
|
||||
echo " and bump the badge in content/index.md (see that repo's CLAUDE.md)"
|
||||
echo " 2. basicmachines.co - Update version in src/components/sections/hero.tsx"
|
||||
echo " 2. basicmemory.com - No version number in the site UI; for a significant"
|
||||
echo " release optionally add a post under src/content/blog/. Skip for patches."
|
||||
echo " 3. MCP Registry - Run: mcp-publisher publish"
|
||||
echo " See: .claude/commands/release/release.md for detailed instructions"
|
||||
|
||||
@@ -585,7 +586,8 @@ beta version:
|
||||
echo "📝 REMINDER: For stable releases, update documentation sites:"
|
||||
echo " 1. docs.basicmemory.com - Add a What's New page under content/2.whats-new/"
|
||||
echo " and bump the badge in content/index.md (see that repo's CLAUDE.md)"
|
||||
echo " 2. basicmachines.co - Update version in src/components/sections/hero.tsx"
|
||||
echo " 2. basicmemory.com - No version number in the site UI; for a significant"
|
||||
echo " release optionally add a post under src/content/blog/. Skip for patches."
|
||||
echo " See: .claude/commands/release/release.md for detailed instructions"
|
||||
|
||||
# List all available recipes
|
||||
|
||||
@@ -18,6 +18,7 @@ from basic_memory.repository.search_index_row import SearchIndexRow
|
||||
from basic_memory.repository.search_repository_base import (
|
||||
SearchRepositoryBase,
|
||||
VectorChunkState,
|
||||
relaxed_query_words,
|
||||
)
|
||||
from basic_memory.repository.metadata_filters import parse_metadata_filters
|
||||
from basic_memory.repository.semantic_errors import SemanticDependenciesMissingError
|
||||
@@ -176,6 +177,14 @@ class PostgresSearchRepository(SearchRepositoryBase):
|
||||
# For non-Boolean queries, prepare single term
|
||||
return self._prepare_single_term(term, is_prefix)
|
||||
|
||||
@staticmethod
|
||||
def _relaxed_tsquery_text(search_text: Optional[str]) -> Optional[str]:
|
||||
"""OR-relaxed tsquery expression for a failed strict query, or None."""
|
||||
words = relaxed_query_words(search_text)
|
||||
if not words:
|
||||
return None
|
||||
return " | ".join(f"{word}:*" for word in words)
|
||||
|
||||
def _prepare_boolean_query(self, query: str) -> str:
|
||||
"""Convert Boolean query to tsquery format.
|
||||
|
||||
@@ -236,7 +245,12 @@ class PostgresSearchRepository(SearchRepositoryBase):
|
||||
|
||||
# Handle multi-word queries
|
||||
if " " in cleaned_term:
|
||||
words = [w for w in cleaned_term.split() if w.strip()]
|
||||
# Strip sentence punctuation from word edges so question-form
|
||||
# queries produce clean lexemes (parity with SQLite FTS5 prep).
|
||||
# The tsquery tokenizer ignores this punctuation anyway; leaving it
|
||||
# in only risks tsquery syntax errors. Interior characters are kept.
|
||||
words = [w.strip("?!.,;") for w in cleaned_term.split()]
|
||||
words = [w for w in words if w]
|
||||
if not words:
|
||||
# All characters were special chars, search won't match anything
|
||||
# Return a safe search term that won't cause syntax errors
|
||||
@@ -249,8 +263,11 @@ class PostgresSearchRepository(SearchRepositoryBase):
|
||||
# Join with AND operator
|
||||
return " & ".join(prepared_words)
|
||||
|
||||
# Single word
|
||||
cleaned_term = cleaned_term.strip()
|
||||
# Single word: strip edge punctuation; guard the now-empty case so a
|
||||
# bare ":*"/"" never reaches tsquery.
|
||||
cleaned_term = cleaned_term.strip().strip("?!.,;")
|
||||
if not cleaned_term:
|
||||
return "NOSPECIALCHARS:*"
|
||||
if is_prefix:
|
||||
return f"{cleaned_term}:*"
|
||||
else:
|
||||
@@ -908,6 +925,7 @@ class PostgresSearchRepository(SearchRepositoryBase):
|
||||
min_similarity: Optional[float] = None,
|
||||
limit: int = 10,
|
||||
offset: int = 0,
|
||||
allow_relaxed: bool = False,
|
||||
) -> List[SearchIndexRow]:
|
||||
"""Search across all indexed content using PostgreSQL tsvector."""
|
||||
# --- Dispatch vector / hybrid modes (shared logic) ---
|
||||
@@ -982,6 +1000,20 @@ class PostgresSearchRepository(SearchRepositoryBase):
|
||||
async with db.scoped_session(self.session_maker) as session:
|
||||
result = await session.execute(text(sql), params)
|
||||
rows = result.fetchall()
|
||||
# Trigger: multi-word natural-language query matched nothing
|
||||
# under the default all-terms-AND tsquery semantics.
|
||||
# Why: questions rarely have every word in one document;
|
||||
# without relaxation the FTS half of hybrid search contributes
|
||||
# zero candidates (parity with the SQLite path).
|
||||
# Outcome: one retry with OR-joined prefix lexemes; ts_rank
|
||||
# still ranks multi-term matches first.
|
||||
relaxed = (
|
||||
self._relaxed_tsquery_text(search_text) if allow_relaxed and not rows else None
|
||||
)
|
||||
if relaxed and params.get("text"):
|
||||
params["text"] = relaxed
|
||||
result = await session.execute(text(sql), params)
|
||||
rows = result.fetchall()
|
||||
except Exception as e:
|
||||
if self._is_tsquery_syntax_error(e):
|
||||
logger.warning(f"tsquery syntax error for search term: {search_text}, error: {e}")
|
||||
|
||||
@@ -40,6 +40,50 @@ BULLET_PATTERN = re.compile(r"^[\-\*]\s+")
|
||||
OVERSIZED_ENTITY_VECTOR_SHARD_SIZE = 256
|
||||
_SQLITE_MAX_PREPARE_WINDOW = 8
|
||||
|
||||
# Interrogative/function words contribute lexical noise when a strict
|
||||
# full-text query is relaxed: "when OR did OR a" matches loud wrong documents
|
||||
# that displace genuine results from the ranking window.
|
||||
RELAXATION_STOPWORDS = frozenset(
|
||||
"a an and are as at be but by did do does for from had has have how i in is it of on "
|
||||
"or that the their they this to was we were what when where which who whom whose why "
|
||||
"will with you your".split()
|
||||
)
|
||||
|
||||
|
||||
def relaxed_query_words(search_text: Optional[str]) -> Optional[list[str]]:
|
||||
"""Content-bearing words for OR-relaxing a strict full-text query.
|
||||
|
||||
Returns None when relaxation must not apply. These eligibility rules match
|
||||
SearchService._is_relaxed_fts_fallback_eligible so the hybrid FTS branch
|
||||
relaxes exactly the same query shapes as the service-level FTS path:
|
||||
|
||||
- empty / quoted / explicit-boolean queries (user intent is not
|
||||
second-guessed);
|
||||
- fewer than three alphanumeric tokens (short queries like "New Feature"
|
||||
over-broaden under OR — and in hybrid the relaxed FTS-only rows normalize
|
||||
to 1.0 and can outrank the vector result the user wanted);
|
||||
- any pure-digit token ("root note 1", "SPEC 16") — identifier-like queries
|
||||
over-broaden and create false positives under OR.
|
||||
"""
|
||||
if not search_text:
|
||||
return None
|
||||
stripped = search_text.strip()
|
||||
if '"' in stripped or any(op in f" {stripped} " for op in (" AND ", " OR ", " NOT ")):
|
||||
return None
|
||||
# Eligibility checks run on raw alphanumeric tokens (parity with the
|
||||
# service), before stopword filtering.
|
||||
tokens = re.findall(r"[A-Za-z0-9]+", stripped.lower())
|
||||
if len(tokens) < 3 or any(token.isdigit() for token in tokens):
|
||||
return None
|
||||
words = [word.strip("?!.,;:") for word in stripped.split()]
|
||||
words = [
|
||||
word
|
||||
for word in words
|
||||
if word and word.isalnum() and word.lower() not in RELAXATION_STOPWORDS
|
||||
]
|
||||
return words or None
|
||||
|
||||
|
||||
# Entity, observation, and relation rows in search_index carry ids from independent
|
||||
# auto-increment sequences, so a bare id is ambiguous across row types. Every map in
|
||||
# the vector/hybrid retrieval path must key rows by (type, id) to avoid collisions.
|
||||
@@ -229,6 +273,7 @@ class SearchRepositoryBase(ABC):
|
||||
min_similarity: Optional[float] = None,
|
||||
limit: int = 10,
|
||||
offset: int = 0,
|
||||
allow_relaxed: bool = False,
|
||||
) -> List[SearchIndexRow]:
|
||||
"""Search across all indexed content.
|
||||
|
||||
@@ -2174,6 +2219,9 @@ class SearchRepositoryBase(ABC):
|
||||
query_start = time.perf_counter()
|
||||
candidate_limit = max(self._semantic_vector_k, (limit + offset) * 10)
|
||||
fts_start = time.perf_counter()
|
||||
# allow_relaxed: question-form queries rarely AND-match, and a dead FTS
|
||||
# branch silently degrades hybrid to vector-only ranking. Fusion plus
|
||||
# bm25 keep relaxed lexical candidates from dominating precision.
|
||||
fts_results = await self.search(
|
||||
search_text=search_text,
|
||||
permalink=permalink,
|
||||
@@ -2187,6 +2235,7 @@ class SearchRepositoryBase(ABC):
|
||||
retrieval_mode=SearchRetrievalMode.FTS,
|
||||
limit=candidate_limit,
|
||||
offset=0,
|
||||
allow_relaxed=True,
|
||||
)
|
||||
fts_ms = (time.perf_counter() - fts_start) * 1000
|
||||
vector_start = time.perf_counter()
|
||||
|
||||
@@ -23,7 +23,10 @@ from basic_memory.models.search import (
|
||||
from basic_memory.repository.embedding_provider import EmbeddingProvider
|
||||
from basic_memory.repository.embedding_provider_factory import create_embedding_provider
|
||||
from basic_memory.repository.search_index_row import SearchIndexRow
|
||||
from basic_memory.repository.search_repository_base import SearchRepositoryBase
|
||||
from basic_memory.repository.search_repository_base import (
|
||||
SearchRepositoryBase,
|
||||
relaxed_query_words,
|
||||
)
|
||||
from basic_memory.repository.metadata_filters import parse_metadata_filters, build_sqlite_json_path
|
||||
from basic_memory.repository.semantic_errors import SemanticDependenciesMissingError
|
||||
from basic_memory.schemas.search import SearchItemType, SearchRetrievalMode
|
||||
@@ -255,6 +258,19 @@ class SQLiteSearchRepository(SearchRepositoryBase):
|
||||
if "*" in term and all(c.isalnum() or c in "*_-" for c in term):
|
||||
return term
|
||||
|
||||
# Natural-language queries arrive with sentence punctuation that FTS5
|
||||
# treats as syntax ("When did Melanie paint a sunrise?"). The tokenizer
|
||||
# ignores this punctuation in the INDEX, so stripping it from word
|
||||
# edges loses nothing — but leaving it forces the whole question into
|
||||
# an exact-phrase match that returns zero rows, silently disabling the
|
||||
# FTS half of hybrid search. Interior characters (hyphens, slashes —
|
||||
# permalinks and paths) are untouched.
|
||||
if " " in term:
|
||||
words = [word.strip("?!.,;:") for word in term.split()]
|
||||
term = " ".join(word for word in words if word)
|
||||
if not term:
|
||||
return ""
|
||||
|
||||
# Characters that can cause FTS5 syntax errors when used as operators
|
||||
# We're more conservative here - only quote when we detect problematic patterns
|
||||
problematic_chars = [
|
||||
@@ -351,6 +367,14 @@ class SQLiteSearchRepository(SearchRepositoryBase):
|
||||
# For non-Boolean queries, use the single term preparation logic
|
||||
return self._prepare_single_term(term, is_prefix)
|
||||
|
||||
@staticmethod
|
||||
def _relaxed_fts_text(search_text: Optional[str]) -> Optional[str]:
|
||||
"""OR-relaxed FTS5 expression for a failed strict query, or None."""
|
||||
words = relaxed_query_words(search_text)
|
||||
if not words:
|
||||
return None
|
||||
return " OR ".join(f"{word}*" for word in words)
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# sqlite-vec extension loading (SQLite-specific)
|
||||
# ------------------------------------------------------------------
|
||||
@@ -953,8 +977,15 @@ class SQLiteSearchRepository(SearchRepositoryBase):
|
||||
min_similarity: Optional[float] = None,
|
||||
limit: int = 10,
|
||||
offset: int = 0,
|
||||
allow_relaxed: bool = False,
|
||||
) -> List[SearchIndexRow]:
|
||||
"""Search across all indexed content using SQLite FTS5."""
|
||||
"""Search across all indexed content using SQLite FTS5.
|
||||
|
||||
``allow_relaxed=True`` retries a zero-result strict multi-word query
|
||||
with OR-joined content terms. Only the hybrid path opts in: its FTS
|
||||
branch otherwise contributes nothing for question-form queries.
|
||||
Service-level FTS searches keep their own conservative fallback.
|
||||
"""
|
||||
# --- Dispatch vector / hybrid modes (shared logic) ---
|
||||
dispatched = await self._dispatch_retrieval_mode(
|
||||
search_text=search_text,
|
||||
@@ -1021,6 +1052,21 @@ class SQLiteSearchRepository(SearchRepositoryBase):
|
||||
async with db.scoped_session(self.session_maker) as session:
|
||||
result = await session.execute(text(sql), params)
|
||||
rows = result.fetchall()
|
||||
# Trigger: multi-word natural-language query matched nothing
|
||||
# under the default all-terms-AND semantics.
|
||||
# Why: questions ("when did X do Y") rarely have every word in
|
||||
# one document; without relaxation the FTS half of hybrid
|
||||
# search contributes zero candidates and ranking degrades to
|
||||
# vector-only.
|
||||
# Outcome: one retry with OR-joined prefix terms; bm25 still
|
||||
# ranks multi-term matches first.
|
||||
relaxed = (
|
||||
self._relaxed_fts_text(search_text) if allow_relaxed and not rows else None
|
||||
)
|
||||
if relaxed and params.get("text"):
|
||||
params["text"] = relaxed
|
||||
result = await session.execute(text(sql), params)
|
||||
rows = result.fetchall()
|
||||
except Exception as e:
|
||||
# Handle FTS5 syntax errors and provide user-friendly feedback
|
||||
if self._is_fts5_syntax_error(e): # pragma: no cover
|
||||
|
||||
@@ -77,6 +77,7 @@ class ConcreteSearchRepo(SearchRepositoryBase):
|
||||
min_similarity: Optional[float] = None,
|
||||
limit: int = 10,
|
||||
offset: int = 0,
|
||||
allow_relaxed: bool = False,
|
||||
) -> list[SearchIndexRow]:
|
||||
return [] # pragma: no cover
|
||||
|
||||
|
||||
@@ -1001,3 +1001,59 @@ async def test_postgres_search_categories_exact_match(session_maker, test_projec
|
||||
# Multiple categories union.
|
||||
multi = await repo.search(categories=["requirement", "decision"])
|
||||
assert {r.id for r in multi} == {70101, 70102}
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_postgres_question_punctuation_and_relaxation(session_maker, test_project):
|
||||
"""Question-form queries must produce clean lexemes and a usable relaxation.
|
||||
|
||||
Parity with SQLite: sentence punctuation previously reached tsquery terms,
|
||||
and a strict all-AND miss had no relaxed retry, silently disabling the FTS
|
||||
half of hybrid search for natural-language questions.
|
||||
"""
|
||||
repo = PostgresSearchRepository(session_maker, project_id=test_project.id)
|
||||
|
||||
# Edge punctuation stripped before lexeme formatting.
|
||||
prepared = repo._prepare_search_term("When did Melanie paint a sunrise?")
|
||||
assert "?" not in prepared
|
||||
assert "sunrise:*" in prepared
|
||||
|
||||
# Relaxation drops stopwords and OR-joins content terms.
|
||||
relaxed = repo._relaxed_tsquery_text("When did Melanie paint a sunrise?")
|
||||
assert relaxed == "Melanie:* | paint:* | sunrise:*"
|
||||
|
||||
# User intent is not second-guessed.
|
||||
assert repo._relaxed_tsquery_text("alpha AND beta") is None
|
||||
assert repo._relaxed_tsquery_text('"exact phrase"') is None
|
||||
assert repo._relaxed_tsquery_text(None) is None
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_postgres_multiword_query_relaxes_on_strict_miss(session_maker, test_project):
|
||||
repo = PostgresSearchRepository(session_maker, project_id=test_project.id)
|
||||
now = datetime.now(timezone.utc)
|
||||
await repo.index_item(
|
||||
SearchIndexRow(
|
||||
project_id=test_project.id,
|
||||
id=77,
|
||||
title="Trip plans",
|
||||
content_stems="melanie painted a sunrise over the lake last year",
|
||||
content_snippet="Melanie painted a sunrise over the lake last year.",
|
||||
permalink="docs/trip-plans",
|
||||
file_path="docs/trip-plans.md",
|
||||
type="entity",
|
||||
metadata={"note_type": "note"},
|
||||
created_at=now,
|
||||
updated_at=now,
|
||||
)
|
||||
)
|
||||
|
||||
# A content word absent from the doc ("hiking") makes the strict
|
||||
# all-terms-AND query miss even after Postgres drops stopwords — without
|
||||
# it, to_tsquery('english', ...) already strips "when/did/a" and matches.
|
||||
strict = await repo.search(search_text="Did Melanie go hiking at sunrise?")
|
||||
assert strict == []
|
||||
|
||||
# The hybrid FTS branch opts in; OR-relaxation surfaces the partial match.
|
||||
results = await repo.search(search_text="Did Melanie go hiking at sunrise?", allow_relaxed=True)
|
||||
assert any(r.id == 77 for r in results)
|
||||
|
||||
@@ -1124,3 +1124,83 @@ async def test_search_categories_exact_match(search_repository, search_entity):
|
||||
# Multiple categories union: both observations come back.
|
||||
multi = await search_repository.search(categories=["requirement", "decision"])
|
||||
assert {r.id for r in multi} == {70001, 70002}
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_question_punctuation_does_not_phrase_quote(search_repository):
|
||||
"""Sentence punctuation must not force exact-phrase matching (#hybrid-fts).
|
||||
|
||||
'When did Melanie paint a sunrise?' previously became the FTS5 phrase
|
||||
'"When did Melanie paint a sunrise?"*' — zero rows for any corpus — which
|
||||
silently disabled the FTS half of hybrid search for question queries.
|
||||
"""
|
||||
prepared = search_repository._prepare_single_term("When did Melanie paint a sunrise?")
|
||||
assert '"' not in prepared
|
||||
# Prefix syntax differs by backend: FTS5 uses '*', tsquery uses ':*'.
|
||||
if is_postgres_backend(search_repository):
|
||||
assert "sunrise:*" in prepared
|
||||
else:
|
||||
assert "sunrise*" in prepared
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_relaxed_query_drops_stopwords(search_repository):
|
||||
"""Relaxation keys on content-bearing terms in each backend's syntax."""
|
||||
if is_postgres_backend(search_repository):
|
||||
relaxed = search_repository._relaxed_tsquery_text("When did Melanie paint a sunrise?")
|
||||
assert relaxed == "Melanie:* | paint:* | sunrise:*"
|
||||
else:
|
||||
relaxed = search_repository._relaxed_fts_text("When did Melanie paint a sunrise?")
|
||||
assert relaxed == "Melanie* OR paint* OR sunrise*"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_relaxed_query_respects_user_intent(search_repository):
|
||||
# Eligibility matches the service-level relaxation (both backends): quoted,
|
||||
# boolean, short (<3 tokens), and numeric-identifier queries are not relaxed.
|
||||
relaxer = (
|
||||
search_repository._relaxed_tsquery_text
|
||||
if is_postgres_backend(search_repository)
|
||||
else search_repository._relaxed_fts_text
|
||||
)
|
||||
assert relaxer("alpha AND beta") is None
|
||||
assert relaxer('"exact phrase"') is None
|
||||
assert relaxer("single") is None # < 3 tokens
|
||||
assert relaxer("New Feature") is None # < 3 tokens (link title)
|
||||
assert relaxer("root note 1") is None # numeric identifier token
|
||||
assert relaxer("SPEC 16 design") is None # numeric token
|
||||
assert relaxer(None) is None
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_multiword_query_relaxes_to_or_when_strict_misses(search_repository, search_entity):
|
||||
"""A question sharing only SOME words with a doc still surfaces it."""
|
||||
from basic_memory.repository.search_index_row import SearchIndexRow
|
||||
from basic_memory.schemas.search import SearchItemType
|
||||
|
||||
row = SearchIndexRow(
|
||||
project_id=search_repository.project_id,
|
||||
id=search_entity.id,
|
||||
type=SearchItemType.ENTITY.value,
|
||||
title="Trip plans",
|
||||
content_snippet="Melanie painted a sunrise over the lake last year.",
|
||||
content_stems="melanie painted a sunrise over the lake last year",
|
||||
permalink=search_entity.permalink,
|
||||
file_path=search_entity.file_path,
|
||||
entity_id=search_entity.id,
|
||||
metadata={"note_type": search_entity.note_type},
|
||||
created_at=search_entity.created_at,
|
||||
updated_at=search_entity.updated_at,
|
||||
)
|
||||
await search_repository.index_item(row)
|
||||
|
||||
# "hiking" is absent from the doc, so strict all-terms-AND misses on both
|
||||
# backends (Postgres's stopword stripping can't rescue it either).
|
||||
strict = await search_repository.search(search_text="Did Melanie go hiking at sunrise?")
|
||||
assert strict == []
|
||||
|
||||
# The hybrid FTS branch opts in; OR-relaxation surfaces the partial match.
|
||||
results = await search_repository.search(
|
||||
search_text="Did Melanie go hiking at sunrise?", allow_relaxed=True
|
||||
)
|
||||
assert any(r.entity_id == search_entity.id for r in results)
|
||||
|
||||
@@ -89,6 +89,7 @@ class _ConcreteRepo(SearchRepositoryBase):
|
||||
min_similarity: float | None = None,
|
||||
limit: int = 10,
|
||||
offset: int = 0,
|
||||
allow_relaxed: bool = False,
|
||||
) -> list[SearchIndexRow]:
|
||||
return []
|
||||
|
||||
|
||||
@@ -62,6 +62,7 @@ class ConcreteSearchRepo(SearchRepositoryBase):
|
||||
min_similarity: float | None = None,
|
||||
limit: int = 10,
|
||||
offset: int = 0,
|
||||
allow_relaxed: bool = False,
|
||||
) -> list[SearchIndexRow]:
|
||||
return [] # pragma: no cover
|
||||
|
||||
|
||||
@@ -66,6 +66,7 @@ class ConcreteSearchRepo(SearchRepositoryBase):
|
||||
min_similarity: Optional[float] = None,
|
||||
limit: int = 10,
|
||||
offset: int = 0,
|
||||
allow_relaxed: bool = False,
|
||||
) -> list[SearchIndexRow]:
|
||||
return [] # pragma: no cover
|
||||
|
||||
|
||||
Reference in New Issue
Block a user