Query Flow
Understand how Rephole processes your semantic code search queries.
Query Pipeline
Question → Embed → Search → Retrieve → Return
When you submit a search query:
- Embed: Your question is converted to a vector using the same embedding model
- Search: ChromaDB performs similarity search against stored vectors
- Retrieve: Matching child chunks identify parent documents
- Return: Full file content is returned with relevant context
How Semantic Search Works
Traditional Code Search
Query: "authenticate"
Results: All files containing the word "authenticate"
Rephole Semantic Search
Query: "How does user login work?"
Results: Authentication functions, login handlers, session management
(even if they don't contain the word "login")
Rephole understands the intent behind your query, not just the keywords.
Search Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
repoId | string | ✅ Yes | - | Repository ID (in URL path) |
prompt | string | ✅ Yes | - | Natural language query |
k | number | No | 5 | Number of results (max: 100) |
meta | object | No | - | Additional metadata filters |
The k Multiplier
Internally, Rephole multiplies k by 3 for child chunk search:
- You request
k=5results - Rephole searches for
15child chunks - Returns
5unique parent documents
This ensures diverse, high-quality results.
Metadata Filtering
Rephole supports custom metadata for organizing and filtering your codebase.
During Ingestion
Tag your repositories with custom metadata:
{
"repoUrl": "https://github.com/org/backend-api.git",
"meta": {
"team": "platform",
"environment": "production",
"version": "2.0"
}
}
During Search
Filter results using metadata in the request body:
# repoId is required in the URL path
curl -X POST http://localhost:3000/queries/search/backend-api \
-H "Content-Type: application/json" \
-d '{
"prompt": "How does caching work?",
"meta": {
"team": "platform"
}
}'
Use Cases
| Use Case | Example Metadata |
|---|---|
| 🏢 Multi-team organizations | {"team": "backend"} |
| 🌍 Multi-environment | {"environment": "production"} |
| 📦 Microservices | {"service": "auth-service"} |
| 🏷️ Project tagging | {"project": "core-api"} |
Filter Logic
Multiple metadata filters are combined with AND logic. All specified filters must match.
Response Format
{
"results": [
{
"id": "src/auth/auth.service.ts",
"content": "import { Injectable } from '@nestjs/common';\n...",
"repoId": "my-repo",
"metadata": {
"team": "backend",
"category": "repository"
}
}
]
}
| Field | Type | Description |
|---|---|---|
id | string | File path |
content | string | Full file content |
repoId | string | Repository identifier |
metadata | object | Custom metadata from ingestion |
Example Queries
Basic Search
curl -X POST http://localhost:3000/queries/search/my-repo \
-H "Content-Type: application/json" \
-d '{"prompt": "How is caching implemented?", "k": 5}'
Search with Metadata Filters
curl -X POST http://localhost:3000/queries/search/backend-api \
-H "Content-Type: application/json" \
-d '{
"prompt": "Database connection pooling",
"k": 5,
"meta": {
"team": "platform",
"environment": "production"
}
}'
Query Tips
Be Specific
❌ "authentication"
✅ "How does JWT token validation work?"
Ask Questions
❌ "database connection"
✅ "Where is the database connection pool configured?"
Describe Intent
❌ "error"
✅ "How are API errors formatted and returned to clients?"
Performance
| Metric | Typical Value |
|---|---|
| Query embedding | ~100ms |
| Vector search | ~50ms |
| Full response | ~200-500ms |
Next Steps
- Explore the API Reference for all endpoints
- Learn about Architecture for scaling