Skip to main content

Query Flow

Understand how Rephole processes your semantic code search queries.

Query Pipeline​

Question → Embed → Search → Retrieve → Return

When you submit a search query:

  1. Embed: Your question is converted to a vector using the same embedding model
  2. Search: ChromaDB performs similarity search against stored vectors
  3. Retrieve: Matching child chunks identify parent documents
  4. Return: Full file content is returned with relevant context

How Semantic Search Works​

Query: "authenticate"
Results: All files containing the word "authenticate"
Query: "How does user login work?"
Results: Authentication functions, login handlers, session management
(even if they don't contain the word "login")

Rephole understands the intent behind your query, not just the keywords.

Search Parameters​

ParameterTypeRequiredDefaultDescription
repoIdstring✅ Yes-Repository ID (in URL path)
promptstring✅ Yes-Natural language query
knumberNo5Number of results (max: 100)
metaobjectNo-Additional metadata filters

The k Multiplier​

Internally, Rephole multiplies k by 3 for child chunk search:

  • You request k=5 results
  • Rephole searches for 15 child chunks
  • Returns 5 unique parent documents

This ensures diverse, high-quality results.

Metadata Filtering​

Rephole supports custom metadata for organizing and filtering your codebase.

During Ingestion​

Tag your repositories with custom metadata:

{
"repoUrl": "https://github.com/org/backend-api.git",
"meta": {
"team": "platform",
"environment": "production",
"version": "2.0"
}
}

Filter results using metadata in the request body:

# repoId is required in the URL path
curl -X POST http://localhost:3000/queries/search/backend-api \
-H "Content-Type: application/json" \
-d '{
"prompt": "How does caching work?",
"meta": {
"team": "platform"
}
}'

Use Cases​

Use CaseExample Metadata
🏢 Multi-team organizations{"team": "backend"}
🌍 Multi-environment{"environment": "production"}
📦 Microservices{"service": "auth-service"}
🏷️ Project tagging{"project": "core-api"}
Filter Logic

Multiple metadata filters are combined with AND logic. All specified filters must match.

Response Format​

{
"results": [
{
"id": "src/auth/auth.service.ts",
"content": "import { Injectable } from '@nestjs/common';\n...",
"repoId": "my-repo",
"metadata": {
"team": "backend",
"category": "repository"
}
}
]
}
FieldTypeDescription
idstringFile path
contentstringFull file content
repoIdstringRepository identifier
metadataobjectCustom metadata from ingestion

Example Queries​

curl -X POST http://localhost:3000/queries/search/my-repo \
-H "Content-Type: application/json" \
-d '{"prompt": "How is caching implemented?", "k": 5}'

Search with Metadata Filters​

curl -X POST http://localhost:3000/queries/search/backend-api \
-H "Content-Type: application/json" \
-d '{
"prompt": "Database connection pooling",
"k": 5,
"meta": {
"team": "platform",
"environment": "production"
}
}'

Query Tips​

Be Specific​

❌ "authentication"
✅ "How does JWT token validation work?"

Ask Questions​

❌ "database connection"
✅ "Where is the database connection pool configured?"

Describe Intent​

❌ "error"
✅ "How are API errors formatted and returned to clients?"

Performance​

MetricTypical Value
Query embedding~100ms
Vector search~50ms
Full response~200-500ms

Next Steps​