- CryptoSentimentCalibrator: 2,086-term weighted lexicon (-100 to +100) * Priority-based span matching (longest-first, no double-count) * Whale phrases ±50, compounds ±30, dot-separated ±20, singles ±10-25 - Calibration logic: lexicon wins on disagreement, amplifies on agreement - CentroidManager: 6 params × 1024-dim built from lexicon via e5-large-v2 - ScoringEngine._refine_with_centroids: fixed attribute access bug - config/centroids/*.npy: padded to 1024-dim (e5-large-v2 output) - lexicon_weights.json: generated unified lexicon - Unit tests: 46/46 core NLP tests pass - Labeling pipeline: 22 samples processed - All 19 critical crypto semantic tests + 6 calibration scenarios pass
32 KiB
Sentiment Engine — Vocabulary / N-gram / Phrase Storage & Scoring Specification
Version: 1.0
Date: 2026-07-16
Scope: Complete inventory of how terms, n-grams, phrases, and their "meaning/score/impact" are stored across the sentiment engine codebase.
1. Executive Summary
The sentiment engine stores vocabulary and scoring signals in six distinct layers, each with different persistence, mutability, and semantics:
| Layer | Storage Format | Mutability | Scope | Primary Use |
|---|---|---|---|---|
| A. Hard-coded Keyword Lists | Python class constants (list[str]) |
Code change + deploy | Crypto-specific sentiment direction (bullish/bearish/whale) | FinBERT calibration override |
| B. Asset Alias Maps | YAML (config/asset_aliases.yaml) |
Config reload / hot-reload | Canonical ticker resolution | Entity extraction → asset_id mapping |
| C. Known Entities Registry | YAML (config/known_entities.yaml) |
Config reload | Asset metadata (chain, contracts, market cap) | Entity enrichment, contract resolution |
| D. Source Credibility Registry | YAML (config/source_credibility.yaml) |
Config reload | Per-source base_credibility + relevance | Credibility scoring, source weighting |
| E. BERT Centroids | NumPy .npy (config/centroids/*.npy) |
Rebuild via encoder | Semantic similarity for 6 scoring parameters | Parameter refinement via embedding similarity |
| F. Labeling Guidelines / Schema | Python enums + docstrings (labeling_pipeline.py) |
Code change | 3-class sentiment, 12-class event, 6-class emotion | Ground-truth label definitions for training |
Critical Observation: There is no single centralized vocabulary store. The system is disjoint by design — each layer serves a different pipeline stage and has its own schema, persistence, and update mechanism.
2. Layer-by-Layer Specification
2.1 Layer A — Hard-coded Keyword Lists (CryptoSentimentCalibrator)
File: src/sentiment_engine/nlp/sentiment_emotion.py
Class: CryptoSentimentCalibrator (lines ~200–1400)
Purpose: Override FinBERT's traditional-finance semantics with crypto-native semantics via keyword matching.
2.1.1 Data Structures
# Four class-level constants — all list[str]
CRYPTO_BULLISH_KEYWORDS: List[str] # ~1,200+ entries
CRYPTO_BEARISH_KEYWORDS: List[str] # ~1,200+ entries
WHALE_BULLISH_PHRASES: List[str] # ~80 entries
WHALE_BEARISH_PHRASES: List[str] # ~120 entries
2.1.2 Entry Format
| Field | Description | Example |
|---|---|---|
| Keyword | Single token or compound phrase with . as space placeholder |
"golden.cross", "whale.accumulation", "surge" |
| Compound phrases | Also duplicated as space-separated strings at list end | "golden cross", "whale accumulation", "all time high" |
Note: The . separator is a convention for internal matching; at runtime, both re.search(r'\b' + re.escape(kw) + r'\b', text_lower) (for single tokens) and simple phrase in text_lower (for whale phrases) are used.
2.1.3 Categories Covered (Bullish)
| Category | Example Keywords |
|---|---|
| Price action | surge, pump, moon, rally, breakout, ath, higher.high |
| Inflows/accumulation | outflow, whale.withdrawal, cold.storage, accumulation, hodl |
| Institutional/ETF | etf, spot.etf, blackrock, fidelity, microstrategy, institutional.adoption |
| Exchange/listing | listing, tier1.listing, binance.listing, coinbase.listing |
| Partnerships/dev | partnership, integration, ecosystem.growth, developer.activity, grant |
| Technical indicators | golden.cross, macd.crossover, rsi.oversold, support.held, 200.day |
| On-chain | whale.accumulation, exchange.outflow, balance.decreasing, staking, hashrate.up |
| DeFi/yield | yield, apy, tvl.growth, protocol.revenue, buyback, token.burn |
| Macro/narrative | halving, supply.shock, inflation.hedge, rate.cut, fed.pivot, risk.on |
| Sentiment/social | fomo, euphoria, optimism, greed, social.dominance, trending |
2.1.4 Categories Covered (Bearish)
| Category | Example Keywords |
|---|---|
| Price action | crash, dump, capitulation, panic, bear.market, lower.high, free.fall |
| Liquidations | liquidation, cascade.liquidation, long.liquidation, margin.call, rekt |
| Hacks/security | hack, exploit, rug, rugpull, stolen, vulnerability, flash.loan.attack |
| Depeg/stablecoin | depeg, stablecoin.depeg, peg.broken, reserve.shortfall, undercollateralized |
| Outflows/selling | inflow, exchange.inflow, balance.increasing, whale.deposit, profit.taking, paper.hands |
| Regulatory | ban, lawsuit, sec.enforcement, crackdown, delist, wells.notice, cease.and.desist |
| Bankruptcy | bankruptcy, insolvency, bank.run, withdrawal.spike, ftx, celcius, terra |
| Technical | death.cross, macd.bearish, rsi.overbought, resistance.held, head.and.shoulders |
| On-chain bearish | whale.selling, exchange.inflow, unstaking, hashrate.down, miner.capitulation |
| DeFi issues | tvl.drop, protocol.exploit, bad.debt, unlock, token.unlock, dilution |
| Macro risk-off | rate.hike, fed.hawkish, tightening, recession, inflation.high, dxy.up, risk.off |
| Sentiment/social | fud, fear, capitulation, despair, anger, narrative.broken, thesis.invalidated |
2.1.5 Whale Action Phrases (Context-Dependent)
| List | Weight | Example Phrases |
|---|---|---|
WHALE_BULLISH_PHRASES |
5× | "whale buys", "whale accumulates", "whale loads", "smart.money.accumulating", "whale.absorbing" |
WHALE_BEARISH_PHRASES |
5× | "whale sells", "whale dumps", "whale distributes", "whale takes profit", "smart.money.selling", "profit taking" |
Weighting: Whale phrases contribute count * 5 to the directional score vs. count * 1 for standard keywords.
2.1.6 Scoring Algorithm (_get_crypto_signal)
def _get_crypto_signal(text: str) -> str:
text_lower = text.lower()
# Whale phrases: simple substring match (higher priority)
whale_bullish = sum(1 for phrase in WHALE_BULLISH_PHRASES if phrase in text_lower)
whale_bearish = sum(1 for phrase in WHALE_BEARISH_PHRASES if phrase in text_lower)
# Standard keywords: word-boundary regex match
bullish_score = sum(1 for kw in CRYPTO_BULLISH_KEYWORDS
if re.search(r'\b' + re.escape(kw) + r'\b', text_lower))
bearish_score = sum(1 for kw in CRYPTO_BEARISH_KEYWORDS
if re.search(r'\b' + re.escape(kw) + r'\b', text_lower))
total_bullish = bullish_score + whale_bullish * 5
total_bearish = bearish_score + whale_bearish * 5
if total_bullish > total_bearish: return "bullish"
elif total_bearish > total_bullish: return "bearish"
return "neutral"
2.1.7 Calibration Logic (calibrate)
The calibrator aggressively flips FinBERT probabilities when crypto keywords disagree:
| Crypto Signal | FinBERT Signal | Action |
|---|---|---|
| bullish | bearish | Force [0.05, neu, 0.95-neu] |
| bearish | bullish | Force [0.95, neu, 0.05] |
| bullish | neutral | Force strong bullish |
| bearish | neutral | Force strong bearish |
| neutral | any | Force neutral (average pos/neg) |
| bullish | bullish | Amplify bullish (+25% of diff) |
| bearish | bearish | Amplify bearish (+50% of diff) |
| any | weak ( | diff |
Key invariant: Crypto keyword signal always wins when FinBERT is uncertain (|pos-neg| < 0.4).
2.1.8 Update Mechanism
- Add/modify: Edit Python source → rebuild container → redeploy
- No hot-reload: Lists are class constants loaded at import time
- Version control: Git history tracks all changes
- Testing:
vocab_test_cases.jsonprovides 200+ regression test cases
2.2 Layer B — Asset Alias Maps
File: config/asset_aliases.yaml
Loaded by: AssetMapper.__init__() → EntityExtractor
Purpose: Map free-text mentions (names, symbols, people) → canonical ticker IDs.
2.2.1 Schema
aliases:
"ALIAS_UPPERCASE": "CANONICAL_TICKER"
# e.g.
"BITCOIN": "BTC"
"ETHEREUM": "ETH"
"VITALIK": "ETH"
"CZ": "BNB"
2.2.2 Entry Types
| Alias Type | Examples | Confidence |
|---|---|---|
| Symbol variants | BTC, XBT → BTC |
0.95 |
| Full names | BITCOIN, ETHEREUM → BTC, ETH |
0.95 |
| Person → asset | VITALIK → ETH, SAYLOR → BTC, ELON → DOGE |
0.7–0.9 |
| Stablecoins | TETHER → USDT, CIRCLE → USDC |
0.95 |
| Memes | SHIBA → SHIB, PEPE → PEPE |
0.95 |
2.2.3 Resolution Logic (AssetMapper.map_ticker)
- Direct alias match (uppercase) → confidence 0.95
- Known entity exact match → confidence 0.9
- Fuzzy match (rapidfuzz, cutoff 85) → confidence 0.8 × similarity
- No match → return as-is, confidence 0.5
2.2.4 Update Mechanism
- Edit YAML → hot-reload on next
AssetMapperinstantiation (no code deploy) - Used by both rule-based extraction (
extract_aliases) and NER post-processing
2.3 Layer C — Known Entities Registry
File: config/known_entities.yaml
Loaded by: AssetMapper._load_known_entities()
Purpose: Rich metadata for canonical assets.
2.3.1 Schema
entities:
BTC:
name: "Bitcoin"
type: "crypto" # crypto | stablecoin | defi | oracle | etc.
chain: "bitcoin"
contracts: [] # empty for native assets
market_cap_rank: 1
ETH:
name: "Ethereum"
type: "crypto"
chain: "ethereum"
contracts: ["0xC02aaA39b223FE8D0A0e5C4F27eAD9083C756Cc2"] # WETH
market_cap_rank: 2
2.3.2 Fields
| Field | Type | Required | Description |
|---|---|---|---|
name |
str | Yes | Human-readable name |
type |
str | Yes | Asset category (crypto, stablecoin, defi, oracle, etc.) |
chain |
str | Yes | Native blockchain |
contracts |
list[str] | No | Contract addresses (for wrapped/bridged versions) |
market_cap_rank |
int | No | Coingecko-style rank |
2.3.3 Usage
- Contract resolution:
AssetMapper.map_contract(address)→ matches againstcontractslist - Fuzzy ticker match:
rapidfuzzagainst entity keys - Entity enrichment:
EntityExtraction.canonical_namepopulated fromname
2.3.4 Update Mechanism
- Edit YAML → hot-reload on next
AssetMapperinstantiation - No code changes required
2.4 Layer D — Source Credibility Registry
File: config/source_credibility.yaml
Loaded by: CatalogueManager._sync_from_config() → CredibilityScorer.load_registry()
Purpose: Per-source base credibility and relevance for weighting signals.
2.4.1 Schema
sources:
- source_id: "rss:coindesk.com"
name: "CoinDesk"
url: "https://www.coindesk.com"
source_type: "news" # news | research | exchange_ann | social | regulatory
base_credibility: 0.85 # 0-1 static prior
relevance: 0.9 # 0-1 crypto relevance
enabled: true
2.4.2 Fields
| Field | Type | Range | Description |
|---|---|---|---|
source_id |
str | — | Unique ID (format: {connector}:{identifier}) |
name |
str | — | Display name |
url |
str | — | Base URL |
source_type |
enum | news, research, exchange_ann, social, regulatory | Category for grouping |
base_credibility |
float | [0,1] | Static prior (updated dynamically at runtime) |
relevance |
float | [0,1] | Domain relevance to crypto markets |
enabled |
bool | — | Whether to ingest from this source |
2.4.3 Runtime Dynamics
- Current credibility (
current_credibility) stored in DuckDB, updated by:- Fetch success/failure rates
- Event outcome feedback (
confirmed+0.02,false_positive-0.05,missed-0.03) - Time decay (half-life 30 days, min 0.1)
- Composite credibility =
current_credibility×relevance× source-type multiplier
2.4.4 Update Mechanism
- YAML edits → hot-reload via
CatalogueManagersync (runs on init + periodic) - Runtime updates persisted to DuckDB (
data/sources.duckdb)
2.5 Layer E — BERT Centroids (Semantic Parameter Scoring)
Files: config/centroids/{fear_state,greed_state,hype_velocity,pub_velocity,pump_score,dump_score}.npy
Managed by: CentroidManager (scoring/centroids.py)
Purpose: Provide semantic "meaning" for 6 scoring parameters via embedding similarity.
2.5.1 Structure
| Parameter | File | Dimension | Description |
|---|---|---|---|
fear_state |
fear_state.npy |
768 (FinBERT) | Fear/panic semantic direction |
greed_state |
greed_state.npy |
768 | Greed/FOMO semantic direction |
hype_velocity |
hype_velocity.npy |
768 | Hype acceleration semantic direction |
pub_velocity |
pub_velocity.npy |
768 | Publication velocity semantic direction |
pump_score |
pump_score.npy |
768 | Pump/manipulation semantic direction |
dump_score |
dump_score.npy |
768 | Dump/crash semantic direction |
2.5.2 Building Process (_build_centroids)
async def _build_centroids(self):
# Current implementation: PLACEHOLDER (random unit vectors)
for param in PARAMETERS:
self._centroids[param] = np.random.randn(768).astype(np.float32)
self._centroids[param] /= np.linalg.norm(self._centroids[param])
TODO (per code comments): Build from keyword lists in SENTIMENT_SPEC_IMPLEMENT_GUIDE.md:
- Collect keyword lists per parameter
- Encode each keyword/sentence via
encoder(e5-large-v2) - Average embeddings → unit vector centroid
- Save to
.npy
2.5.3 Scoring Usage (_refine_with_centroids)
embedding = self._get_text_embedding(combined_text) # e5-large-v2
similarity = centroid_manager.compute_similarity(embedding, param_name)
centroid_score = (similarity + 1.0) / 2.0 # map [-1,1] → [0,1]
params[param_name] = 0.7 * current_value + 0.3 * centroid_score
Weight: 30% centroid similarity, 70% signal-processor value.
2.5.4 Update Mechanism
- Current: Placeholder — random vectors on first init if
.npymissing - Production: Re-run
_build_centroidswith trained encoder → overwrite.npyfiles - No hot-reload: Centroids loaded once at
ScoringEngine.initialize()
2.6 Layer F — Labeling Schema & Guidelines
File: labeling_pipeline.py (lines 1–400+)
Purpose: Define ground-truth label space for supervised training/annotation.
2.6.1 Sentiment Labels (3-class)
| Label | Value | Description |
|---|---|---|
BEARISH |
0 | Explicit negative price expectation |
BULLISH |
1 | Explicit positive price expectation |
NEUTRAL |
2 | No clear directional bias |
Guidelines (from LABELING_GUIDELINES):
| Label | Explicit Keywords | Technical | Fundamental | Emoji |
|---|---|---|---|---|
| BULLISH | "moon", "pump", "accumulate", "to $100k" | "golden cross", "breakout", "higher highs" | "institutional adoption", "ETF approval", "whale accumulation" | 🚀 📈 💎 🙌 🌙 |
| BEARISH | "crash incoming", "dump it", "top is in" | "death cross", "breakdown", "lower high" | "SEC lawsuit", "exchange hack", "regulation ban" | 📉 😭 💀 🩸 🧻 |
| NEUTRAL | "BTC at $50k", "market consolidating" | — | — | — |
2.6.2 Event Types (12-class)
| Index | Label | Description |
|---|---|---|
| 0 | listing |
New exchange listing, token debut |
| 1 | delisting |
Removal from exchange |
| 2 | hack |
Exploit, drain, theft, vulnerability |
| 3 | regulatory |
SEC, CFTC, lawsuits, regulation |
| 4 | governance |
DAO votes, proposals, treasury |
| 5 | upgrade |
Hard fork, mainnet, protocol upgrade |
| 6 | partnership |
Integration, collaboration, alliance |
| 7 | earnings |
Revenue, profit, financial results |
| 8 | macro |
Fed, rates, CPI, GDP, employment |
| 9 | liquidation |
Margin calls, cascade liquidations |
| 10 | whale |
Large transfers, accumulation, distribution |
| 11 | manipulation |
Wash trading, spoofing, pump & dump |
2.6.3 Emotion Types (6-class)
| Label | Keywords |
|---|---|
joy |
moon, pump, breakout, profit, gains, success |
fear |
crash, hack, panic, worry, risk |
anger |
rug, scam, fraud, manipulation, unfair |
greed |
fomo, ape, yolo, leverage, accumulation |
sadness |
loss, rekt, down, bear, pain |
neutral |
sideways, stable, consolidating, range |
2.6.4 Entity Types
TICKER, CONTRACT, PROTOCOL, EXCHANGE, PERSON, CHAIN, ORG
2.6.5 Real Events (Ground Truth)
REAL_EVENTS list in labeling_pipeline.py — 50+ manually labeled examples with text, label_id, event_type.
2.6.6 Update Mechanism
- Edit Python enums/docstrings → rebuild
REAL_EVENTSextended manually for regression testing- Used by
LabelingPipelineRunnerfor automated annotation
3. Pipeline Flow — How Vocabulary Flows Through the System
┌─────────────────────────────────────────────────────────────────────────────────┐
│ SENTIMENT ENGINE VOCABULARY FLOW │
└─────────────────────────────────────────────────────────────────────────────────┘
RAW TEXT INPUT
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ ENTITY EXTRACTION (EntityExtractor) │
│ • Ticker regex: \$?[A-Za-z]{2,10}\b │
│ • Contract regex: 0x[a-fA-F0-9]{40} | base58 │
│ • Alias lookup: Layer B (asset_aliases.yaml) + Layer C (known_entities) │
│ • NER (spaCy): ORG, PRODUCT, GPE, PERSON → fuzzy map to tickers │
│ Output: List[EntityExtraction{asset_id, mention_span, confidence, type}] │
└─────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ SENTIMENT & EMOTION ANALYSIS (SentimentEmotionAnalyzer) │
│ • FinBERT (ONNX/PyTorch/Mock) → [neg, neu, pos] probs │
│ • CryptoSentimentCalibrator.calibrate(text, probs) ← LAYER A KEYWORDS │
│ - _get_crypto_signal() uses CRYPTO_BULLISH/BEARISH_KEYWORDS │
│ - WHALE_*_PHRASES weighted 5× │
│ - Word-boundary regex for standard, substring for whale phrases │
│ • Emotion model (DistilRoBERTa) → 6-class emotions │
│ • Heuristic fallback if models unavailable │
│ Output: SentimentScores(polarity, confidence, pos/neg/neu), EmotionScores │
└─────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ EVENT CLASSIFICATION (EventClassifier) │
│ • BERT classifier → 12-class event type │
│ • Uses Layer F label schema (EVENT_LABELS) │
└─────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ CREDIBILITY SCORING (CredibilityScorer) │
│ • Source base_credibility from Layer D (source_credibility.yaml) │
│ • Cross-source corroboration (in-memory cache) │
│ • Temporal decay (half-life 30 days) │
│ Output: CredibilityScore(composite, components) │
└─────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ SIGNAL PROCESSING (SignalProcessor) │
│ • fear_state = f(neg_sentiment, fear_emotion, event_fear) │
│ • greed_state = f(pos_sentiment, greed_emotion, event_greed) │
│ • pump_score = f(greed, joy, pos_events, intensity) │
│ • dump_score = f(fear, anger, neg_events, intensity) │
│ • VelocityComputer → hype_velocity, pub_velocity │
│ • TemporalDecay (Layer E scoring.halflife_minutes) │
│ Output: AssetSentiment per asset │
└─────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ CENTROID REFINEMENT (ScoringEngine._refine_with_centroids) ← LAYER E │
│ • Embed combined entity+event text via e5-large-v2 │
│ • Cosine similarity to 6 parameter centroids (Layer E .npy files) │
│ • Blend: 70% signal, 30% centroid │
└─────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ AGGREGATION (Aggregator) │
│ • Asset → Industry (Layer C asset_industry_map.yaml) │
│ • Industry → Market │
│ • Decay at each level (asset 30m, industry 60m, market 120m half-life) │
│ Output: SentimentOutput(market, industries, assets) │
└─────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ TRADING INTEGRATION │
│ • ACB signals: market_sentiment_state, fear_state, greed_state, │
│ hype_velocity, aggregate_pump_risk │
│ • Book health veto: pump_score > 75 │
│ • AlphaExitV7: dump_score > 70, fear_state > 80 │
└─────────────────────────────────────────────────────────────────────────────┘
4. Consistency & Centralization Analysis
4.1 Current State: DISJOINT
| Aspect | Status | Detail |
|---|---|---|
| Single source of truth | ❌ No | 6 independent stores with different schemas |
| Unified ID space | ❌ No | Keywords (strings), aliases (ticker→ticker), entities (ticker→metadata), sources (source_id), centroids (param name), labels (enum values) |
| Versioning | Partial | Git for code (Layer A, F), file mtime for YAML (B, C, D), file mtime for .npy (E) |
| Audit trail | Partial | Git for code; DuckDB audit log for source credibility (D); none for centroids (E) |
| Hot-reload | Mixed | YAML (B, C, D): yes; Python constants (A, F): no; .npy (E): no |
| Validation | Minimal | vocab_test_cases.json tests Layer A only; no cross-layer validation |
4.2 Duplication & Drift Risks
| Risk | Location | Example |
|---|---|---|
| Keyword ↔ Label drift | Layer A vs Layer F | CRYPTO_BULLISH_KEYWORDS contains "moon" but LABELING_GUIDELINES lists "moon" under BULLISH emoji — consistent now, but no enforcement |
| Alias ↔ Entity drift | Layer B vs Layer C | asset_aliases.yaml has "VITALIK" → "ETH"; known_entities.yaml has ETH entry — if one updated without other, resolution breaks |
| Centroid ↔ Keyword drift | Layer E vs Layer A | Centroids built from keywords (TODO) but currently random; if keywords change, centroids stale |
| Source credibility ↔ Event outcome | Layer D vs Labeling | false_positive event outcome adjusts credibility but event labels from Layer F — no automated loop |
5. Recommendations for Centralization
5.1 Immediate (Low Effort)
-
Single Vocabulary Registry — Create
config/vocabulary.yamlwith:sentiment_keywords: bullish: [...] bearish: [...] whale_bullish: [...] whale_bearish: [...] asset_aliases: {...} # merge Layer B known_entities: {...} # merge Layer C source_credibility: [...] # merge Layer D labeling_schema: # mirror Layer F sentiment: [BEARISH, BULLISH, NEUTRAL] events: [...] emotions: [...] -
Runtime Loader —
VocabularyRegistryclass loading YAML +.npycentroids, exposing typed accessors. -
Validation Tests — Cross-layer consistency checks:
- Every alias target exists in known_entities
- Every whale phrase keyword appears in corresponding bullish/bearish list
- Centroid rebuild script reads from
vocabulary.yamlkeyword lists
5.2 Medium Term
-
Centroid Auto-Rebuild — On vocabulary change, trigger centroid recomputation via encoder.
-
Provenance Tracking — Add
source: "keyword_list" | "centroid" | "heuristic"to every score component. -
A/B Testing Framework — Compare keyword-only vs. centroid-only vs. blended scoring.
5.3 Long Term
-
Learned Vocabulary — Replace hard-coded lists with learned token importance (attention weights, SHAP values) from fine-tuned model.
-
Semantic Versioning —
vocabulary.yamlwithversion: "2.1.0", migration scripts for schema changes.
6. File Inventory (Absolute Paths)
| Layer | File | Lines | Size | Last Modified |
|---|---|---|---|---|
| A | /mnt/dolphinng5_predict/sentiment_engine/src/sentiment_engine/nlp/sentiment_emotion.py |
~1,776 | ~68 KB | 2026-07-xx |
| B | /mnt/dolphinng5_predict/sentiment_engine/config/asset_aliases.yaml |
~60 | 1.1 KB | 2026-07-xx |
| C | /mnt/dolphinng5_predict/sentiment_engine/config/known_entities.yaml |
~55 | 1.8 KB | 2026-07-xx |
| D | /mnt/dolphinng5_predict/sentiment_engine/config/source_credibility.yaml |
~70 | 2.9 KB | 2026-07-xx |
| E | /mnt/dolphinng5_predict/sentiment_engine/config/centroids/*.npy (6 files) |
— | 3.1 KB each | 2026-07-xx |
| F | /mnt/dolphinng5_predict/sentiment_engine/labeling_pipeline.py |
~1,000+ | ~48 KB | 2026-07-xx |
| Config | /mnt/dolphinng5_predict/sentiment_engine/config/settings.yaml |
~180 | 7.8 KB | 2026-07-xx |
| Test | /mnt/dolphinng5_predict/sentiment_engine/vocab_test_cases.json |
~2,000 | 47 KB | 2026-07-xx |
7. Keyword Counts (Layer A)
| List | Count (approx) | Unique Stems |
|---|---|---|
CRYPTO_BULLISH_KEYWORDS |
1,200+ | ~400 |
CRYPTO_BEARISH_KEYWORDS |
1,200+ | ~400 |
WHALE_BULLISH_PHRASES |
80 | 80 |
WHALE_BEARISH_PHRASES |
120 | 120 |
| Total | ~2,600 | ~1,000 |
Note: High duplication in lists (many variants: "surge", "surges", "surged", "surgeing", "surgeing").
8. Test Coverage (Layer A)
File: vocab_test_cases.json — 200+ test cases
Coverage: Basic positive/negative, whale phrases, compound phrases, edge cases
Run: pytest tests/test_crypto_sentiment_calibrator.py (if exists) or manual via labeling_pipeline.py
9. Open Questions / TODOs
-
Centroid building —
_build_centroids()currently uses random vectors. Implement keyword-driven centroid construction perSENTIMENT_SPEC_IMPLEMENT_GUIDE.md. -
Whale phrase matching — Currently uses simple substring (
phrase in text_lower). Should use word-boundary regex for consistency with standard keywords. -
Compound phrase deduplication — Lists contain both
"golden.cross"and"golden cross". Normalize to single representation. -
Multi-word n-gram storage — No explicit n-gram store beyond compound phrases in keyword lists. Consider adding n-gram frequency tracking from corpus.
-
Language support — Only English (
supported_languages: ["en"]). Keyword lists are English-only. -
Dynamic keyword weighting — All keywords equal weight (1). Could learn weights from labeled data.
10. Appendices
10.1 Full Keyword List Excerpt (Layer A)
See sentiment_emotion.py lines 200–1400 for complete lists.
10.2 Centroid Rebuild Procedure (When Implemented)
# 1. Update vocabulary.yaml with new keywords
# 2. Run rebuild script
python -m sentiment_engine.scripts.rebuild_centroids
# 3. Verify .npy files updated
# 4. Restart scoring engine
10.3 Hot-Reload Procedures
| Layer | Command |
|---|---|
| B, C, D | POST /admin/reload-catalogue (if API exposed) or restart CatalogueManager |
| E | Restart ScoringEngine (no hot-reload) |
| A, F | Full container rebuild + deploy |
End of Specification