Ensemble Composition with Reciprocal Rank Fusion¶
Overview¶
Ensemble composition allows the system to query multiple backend profiles (embedding models) in parallel and intelligently fuse their results using Reciprocal Rank Fusion (RRF). This approach leverages the complementary strengths of different embedding models to improve search quality, particularly on complex queries.
Architecture¶
Components¶
flowchart TB
subgraph Agent["<span style='color:#000'><b>SearchAgent</b></span>"]
Query["<span style='color:#000'>Query</span>"]
end
subgraph Single["<span style='color:#000'><b>Single Profile Mode</b></span>"]
Encode["<span style='color:#000'>Encode</span>"]
Search["<span style='color:#000'>Search</span>"]
Results1["<span style='color:#000'>Results</span>"]
end
subgraph Ensemble["<span style='color:#000'><b>Ensemble Mode</b></span>"]
P1["<span style='color:#000'>Profile 1<br/>ColPali</span>"]
P2["<span style='color:#000'>Profile 2<br/>X-CLIP</span>"]
P3["<span style='color:#000'>Profile 3<br/>Qwen</span>"]
end
RRF["<span style='color:#000'>RRF Fusion</span>"]
Fused["<span style='color:#000'>Fused Results</span>"]
Query --> Encode
Encode --> Search
Search --> Results1
Query --> P1
Query --> P2
Query --> P3
P1 --> RRF
P2 --> RRF
P3 --> RRF
RRF --> Fused
style Agent fill:#ce93d8,stroke:#7b1fa2,color:#000
style Single fill:#90caf9,stroke:#1565c0,color:#000
style Ensemble fill:#a5d6a7,stroke:#388e3c,color:#000
style Query fill:#ba68c8,stroke:#7b1fa2,color:#000
style Encode fill:#64b5f6,stroke:#1565c0,color:#000
style Search fill:#64b5f6,stroke:#1565c0,color:#000
style Results1 fill:#64b5f6,stroke:#1565c0,color:#000
style P1 fill:#81c784,stroke:#388e3c,color:#000
style P2 fill:#81c784,stroke:#388e3c,color:#000
style P3 fill:#81c784,stroke:#388e3c,color:#000
style RRF fill:#ffcc80,stroke:#ef6c00,color:#000
style Fused fill:#a5d6a7,stroke:#388e3c,color:#000 Reciprocal Rank Fusion (RRF)¶
Algorithm¶
RRF is a simple yet effective rank aggregation method that combines rankings from multiple sources without requiring score calibration.
Formula:
Where:
doc: Document/result being scoredk: Constant (default: 60) - controls the weight of top-ranked documentsrank_in_profile: Rank of document in a specific profile's results (0-indexed)
Example¶
Given 3 profiles ranking a document differently (0-indexed):
- Profile 1 (ColPali): rank = 1 (2nd position)
- Profile 2 (X-CLIP): rank = 4 (5th position)
- Profile 3 (Qwen): rank = 0 (1st position)
RRF score = 1/(60+1) + 1/(60+4) + 1/(60+0) = 1/61 + 1/64 + 1/60 = 0.0164 + 0.0156 + 0.0167 = 0.0487
Properties¶
- Score Normalization: Scores are bounded (0, 1/k), making them comparable across profiles
- Rank-Based: Uses only rank information, not raw scores (handles score distribution differences)
- Top-Heavy: Higher-ranked documents get disproportionately more weight
- Unsupervised: No training required, works out-of-the-box
Parameter k¶
The constant k controls the influence of ranking position:
- Lower k (e.g., 30): More weight to top-ranked documents, aggressive fusion
- Higher k (e.g., 100): More equal treatment across ranks, conservative fusion
- Default k=60: Balanced - empirically proven effective across IR tasks
Implementation¶
SearchAgent Ensemble Methods¶
1. Ensemble Search Method¶
async def _search_ensemble(
self,
query: str,
*,
tenant_id: str,
profiles: List[str],
modality: str = "video",
top_k: int = 10,
rrf_k: int = 60,
**kwargs,
) -> List[Dict[str, Any]]:
"""
Execute parallel search across multiple profiles and fuse with RRF.
Args:
query: Text search query
tenant_id: Tenant identifier (per-request, keyword-only)
profiles: List of profile names to query
modality: Content modality to search
top_k: Number of final results to return
rrf_k: RRF constant for fusion
**kwargs: Additional search parameters
Returns:
Fused results from all profiles
"""
2. Parallel Execution¶
The ensemble search implementation uses asyncio.gather to execute searches across all profiles concurrently. Encoding is deduplicated by embedding model first — profiles that share a model (e.g. two ColPali variants) encode the query once and fan the embedding back out, since encoder construction can be a full model load:
# Inside _search_ensemble method:
# Map each profile to its embedding-model "unit"; profiles sharing a
# model share one encode call.
unit_by_profile = {p: _encode_unit(p) for p in profiles}
representative = {}
for profile_name, unit in unit_by_profile.items():
representative.setdefault(unit, profile_name)
# Encode once per distinct unit, in parallel (off the event loop)
unit_embeddings = dict(
await asyncio.gather(
*(encode_unit(u, rep) for u, rep in representative.items())
)
)
# Fan the shared embeddings back out to every profile
valid_embeddings = {
profile: unit_embeddings[unit]
for profile, unit in unit_by_profile.items()
if unit_embeddings.get(unit) is not None
}
# Execute searches in parallel using shared thread pool
with concurrent.futures.ThreadPoolExecutor(max_workers=len(valid_embeddings)) as executor:
search_tasks = [
search_profile(profile, embeddings, executor)
for profile, embeddings in valid_embeddings.items()
]
profile_results_list = await asyncio.gather(*search_tasks)
3. RRF Fusion¶
def _fuse_results_rrf(
self,
profile_results: Dict[str, List[Dict[str, Any]]],
k: int = 60,
top_k: int = 10
) -> List[Dict[str, Any]]:
"""
Fuse results from multiple profiles using Reciprocal Rank Fusion (RRF).
Formula: score(doc) = Σ_profiles (1 / (k + rank_in_profile))
Args:
profile_results: Dict mapping profile names to their result lists
k: RRF constant (default 60, typical range: 20-100)
top_k: Number of final results to return
Returns:
Fused and re-ranked results
"""
# Accumulate RRF scores by document ID
doc_scores = {} # doc_id -> {score, result_data}
# Calculate RRF scores
for profile_name, results in profile_results.items():
for rank, result in enumerate(results):
doc_id = result["id"]
rrf_score = 1.0 / (k + rank)
if doc_id not in doc_scores:
doc_scores[doc_id] = {
"score": 0.0,
"result": result,
"profile_ranks": {},
"profile_scores": {},
}
doc_scores[doc_id]["score"] += rrf_score
doc_scores[doc_id]["profile_ranks"][profile_name] = rank
doc_scores[doc_id]["profile_scores"][profile_name] = result.get("score", 0.0)
# Sort by RRF score
fused_results = []
for doc_id, doc_data in doc_scores.items():
result = doc_data["result"].copy()
result["rrf_score"] = doc_data["score"]
result["profile_ranks"] = doc_data["profile_ranks"]
result["profile_scores"] = doc_data["profile_scores"]
result["num_profiles"] = len(doc_data["profile_ranks"])
fused_results.append(result)
# Sort by RRF score (descending)
fused_results.sort(key=lambda x: x["rrf_score"], reverse=True)
return fused_results[:top_k]
When to Use Ensemble¶
Use Cases¶
Ensemble is beneficial for:
- Complex queries: Multiple semantic aspects (e.g., "show me video of robots playing soccer in tournaments")
- Ambiguous queries: Queries that could match multiple interpretations
- Multi-modal content: Content with both visual and textual components
- Recall-critical tasks: When missing relevant documents is costly
Single profile is sufficient for:
- Simple keyword queries: Direct term matches (e.g., "cat videos")
- Latency-critical applications: When <100ms difference matters
- High-confidence queries: When one profile clearly dominates
Automatic Selection¶
ProfileSelectionAgent can help select the best profile for a query:
class ProfileSelectionSignature(dspy.Signature):
"""Select optimal backend profile based on query analysis"""
query: str = dspy.InputField(desc="User query to analyze")
available_profiles: str = dspy.InputField(
desc="Comma-separated list of available profiles"
)
selected_profile: str = dspy.OutputField(desc="Best matching profile name")
confidence: str = dspy.OutputField(desc="Confidence score 0.0-1.0")
reasoning: str = dspy.OutputField(desc="Explanation for profile selection")
query_intent: ProfileQueryIntent = dspy.OutputField(
desc="Detected intent: multi_modal_search, video_search, image_search, text_search, audio_search, document_search, relationship_aware_search, ensemble_search, code_search, wiki_search"
)
modality: str = dspy.OutputField(
desc="Target modality: audio, code, document, image, text, video, wiki"
)
complexity: Literal["simple", "medium", "complex"] = dspy.OutputField(
desc="Query complexity: simple, medium, complex"
)
Note: ProfileSelectionAgent currently selects a single profile. When the orchestrator chains profile_selection_agent ahead of search_agent, _merge_enrichment() wraps the winning selected_profile as a single-element SearchInput.profiles list (a single-profile override, not ensemble). For ensemble search, explicitly provide multiple profiles via the profiles parameter in SearchInput.
Performance Characteristics¶
Latency¶
| Configuration | Typical Latency | Notes |
|---|---|---|
| Single profile | 400-600ms | Baseline |
| Ensemble (2 profiles) | 500-700ms | +100-150ms overhead |
| Ensemble (3 profiles) | 550-750ms | +150-200ms overhead |
| RRF fusion | 5-10ms | Negligible |
Key insight: Parallel execution keeps ensemble latency close to single-profile latency (not 2x or 3x).
Quality Improvements¶
Ensemble search typically provides significant quality improvements for complex queries. Expected improvements based on information retrieval research:
| Metric | Single Best Profile | Ensemble (3 profiles) | Expected Improvement |
|---|---|---|---|
| NDCG@10 | Baseline | Higher | +10-20% typical |
| MRR | Baseline | Higher | +8-15% typical |
| Recall@20 | Baseline | Higher | +15-25% typical |
Note: Actual improvements vary by query complexity, profile diversity, and content characteristics. Complex queries with multiple semantic aspects tend to benefit most.
Complex queries = queries with >3 entities, >2 relationships, or multi-aspect semantics
Resource Usage¶
- Network: 2-3x connections (parallel requests to Vespa)
- Memory: O(n_profiles × n_results) ~ 5-10KB for typical case
- CPU: Minimal (RRF is O(n) and runs in <10ms)
Configuration¶
Profile Configuration¶
Profiles are defined in config.json:
{
"backend": {
"type": "vespa",
"profiles": {
"video_colpali_smol500_mv_frame": {
"type": "video",
"description": "Frame-based ColPali using TomoroAI/tomoro-colqwen3-embed-4b for 320-dimensional per-patch visual embeddings",
"embedding_model": "TomoroAI/tomoro-colqwen3-embed-4b",
"embedding_type": "multi_vector",
"schema_config": {
"embedding_dim": 320
}
},
"video_xclip_sv_chunk_6s": {
"type": "video",
"description": "X-CLIP base model for 30-second chunk embeddings with 768-dim global representations",
"embedding_model": "microsoft/xclip-large-patch14",
"embedding_type": "multi_vector",
"schema_config": {
"embedding_dim": 768
}
},
"video_colqwen_omni_mv_chunk_30s": {
"type": "video",
"description": "ColQwen3 visual document retrieval served by the Cogniverse ColPali service. 320-dim per-patch multi-vector embeddings.",
"embedding_model": "TomoroAI/tomoro-colqwen3-embed-4b",
"embedding_type": "multi_vector",
"schema_config": {
"embedding_dim": 320
}
}
}
}
}
Ensemble Configuration¶
Ensemble search is configured via SearchInput parameters:
from cogniverse_agents.search_agent import SearchInput
# Configure ensemble search
search_input = SearchInput(
query="robots playing soccer",
tenant_id="acme_corp",
modality="video",
profiles=["video_colpali_smol500_mv_frame", "video_xclip_sv_chunk_6s"],
top_k=10,
rrf_k=60, # RRF constant for fusion
)
Best Practices¶
1. Profile Diversity¶
Choose profiles with complementary strengths:
- ✅ Good: ColPali (visual) + X-CLIP (temporal) + Qwen (cross-modal)
- ❌ Poor: ColPali + ColPali-Large + ColPali-XL (redundant)
2. Limit Ensemble Size¶
Use 2-3 profiles maximum:
- More profiles = diminishing returns
- Complexity increases: O(n_profiles × top_k) for RRF fusion
- Network overhead grows linearly
3. Profile Ordering¶
Order profiles by expected relevance (for early stopping):
# ProfileSelectionAgent should rank profiles
selected_profiles = ["colpali", "xclip", "qwen"] # Best first
4. Conditional Ensemble¶
Don't always use ensemble:
if query_complexity > threshold or confidence < threshold:
use_ensemble = True
else:
use_single_profile = True
5. Monitoring¶
Track ensemble effectiveness:
metrics = {
"ensemble_usage_rate": 0.35, # 35% of queries use ensemble
"quality_improvement": 0.15, # +15% NDCG
"latency_overhead": 150, # +150ms average
"profile_agreement": 0.42, # 42% result overlap
}
Troubleshooting¶
Low Quality Improvement¶
Symptom: Ensemble doesn't improve over single best profile
Possible causes:
- Profiles too similar (high overlap)
- Query too simple (single aspect)
- RRF k value suboptimal
Solutions:
- Choose more diverse profiles
- Use single profile for simple queries
- Tune k parameter (try 30, 60, 100)
High Latency¶
Symptom: Ensemble takes >1s
Possible causes:
- Sequential execution (bug)
- Slow profiles in ensemble
- Network issues
Solutions:
- Verify parallel execution (check logs)
- Remove slow profiles from ensemble
- Increase connection pool size
No Result Overlap¶
Symptom: RRF produces sparse results (few documents ranked by multiple profiles)
Possible causes:
- Profiles searching different indices
- Profiles optimized for different modalities
- Query mismatch
Solutions:
- Ensure all profiles search same content
- Check profile compatibility
- Log profile results for debugging
Multi-Query Fusion¶
Multi-query fusion is a complementary technique to ensemble search. While ensemble varies the profile (embedding model) with a fixed query, multi-query fusion varies the query (via ComposableQueryAnalysisModule LLM-generated variants) against a single profile.
Comparison¶
| Ensemble (Multi-Profile) | Multi-Query Fusion | |
|---|---|---|
| What varies | Profile (embedding model) | Query (rewritten variants) |
| Entry path | _process_impl() → _search_ensemble() | _process_impl() → _search_multi_query_fusion() |
| Input | SearchInput.profiles | SearchInput.query_variants (enrichment field) |
| Fusion | RRF across profiles | RRF across query variants |
| Config | SearchInput.profiles list | SearchInput.rrf_k + SearchInput.query_variants |
How It Works¶
OrchestratorAgentcallsComposableQueryAnalysisModule.forward(query)which extracts entities, infers relationships, enhances the query, and generates query variants in a single composable step- Each variant is encoded and searched in parallel against the same profile
- Results are fused with
_fuse_results_rrf()(same algorithm as ensemble)
Configuration¶
Multi-query fusion has no dedicated config block. The RRF constant rides on the request as SearchInput.rrf_k (default 60) — the same field ensemble fusion uses — and the orchestrator also reads it from context.routing_metadata (rrf_k, default 60). The variants themselves arrive on SearchInput.query_variants.
The composable module's path selection is configured on the relationship router (dspy_relationship_router): - entity_confidence_threshold (default: 0.6) — GLiNER confidence threshold for Path A vs Path B - min_entities_for_fast_path (default: 1) — minimum entities required for Path A
Mutual Exclusivity¶
Ensemble and multi-query fusion use disjoint activation conditions — ensemble activates when SearchInput.profiles has multiple entries; multi-query fusion activates when SearchInput.query_variants has more than one entry (populated by the orchestrator's _merge_enrichment from QueryEnhancementAgent). Both fields are present on SearchInput but only one drives execution per request.
Future Enhancements¶
1. Learned Fusion¶
Replace RRF with learned fusion model:
class LearnedFusion(dspy.Module):
def __init__(self):
self.fusion = dspy.ChainOfThought(FusionSignature)
def forward(self, profile_results, query_features):
# LLM-based intelligent fusion
return self.fusion(results=profile_results, features=query_features)
2. Adaptive k¶
Learn optimal k per query type:
3. Profile Pruning¶
Dynamically select subset of profiles during fusion:
# If profile contributes <5% unique results, remove it
active_profiles = prune_low_contribution_profiles(profile_results)
4. Cross-Modal Reranking¶
Rerank fused results using cross-modal similarity:
reranked = cross_modal_reranker(
fused_results,
query_text=query,
query_embedding=query_emb,
multimodal_features=extracted_features
)
References¶
-
RRF Original Paper: Cormack, G. V., Clarke, C. L., & Buettcher, S. (2009). "Reciprocal rank fusion outperforms condorcet and individual rank learning methods." SIGIR.
-
Multi-Vector Search: Khattab, O., & Zaharia, M. (2020). "ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT." SIGIR.
-
Ensemble Learning in IR: Fox, E. A., & Shaw, J. A. (1994). "Combination of multiple searches." TREC.
See Also¶
- Multi-Agent Interactions - Overall architecture
- Dynamic Profiles - Profile management
- Profile Management - User documentation