# Sentiment Engine — Vocabulary / N-gram / Phrase Storage & Scoring Specification **Version:** 1.0 **Date:** 2026-07-16 **Scope:** Complete inventory of how terms, n-grams, phrases, and their "meaning/score/impact" are stored across the sentiment engine codebase. --- ## 1. Executive Summary The sentiment engine stores vocabulary and scoring signals in **six distinct layers**, each with different persistence, mutability, and semantics: | Layer | Storage Format | Mutability | Scope | Primary Use | |-------|---------------|------------|-------|-------------| | **A. Hard-coded Keyword Lists** | Python class constants (`list[str]`) | Code change + deploy | Crypto-specific sentiment direction (bullish/bearish/whale) | FinBERT calibration override | | **B. Asset Alias Maps** | YAML (`config/asset_aliases.yaml`) | Config reload / hot-reload | Canonical ticker resolution | Entity extraction → asset_id mapping | | **C. Known Entities Registry** | YAML (`config/known_entities.yaml`) | Config reload | Asset metadata (chain, contracts, market cap) | Entity enrichment, contract resolution | | **D. Source Credibility Registry** | YAML (`config/source_credibility.yaml`) | Config reload | Per-source base_credibility + relevance | Credibility scoring, source weighting | | **E. BERT Centroids** | NumPy `.npy` (`config/centroids/*.npy`) | Rebuild via encoder | Semantic similarity for 6 scoring parameters | Parameter refinement via embedding similarity | | **F. Labeling Guidelines / Schema** | Python enums + docstrings (`labeling_pipeline.py`) | Code change | 3-class sentiment, 12-class event, 6-class emotion | Ground-truth label definitions for training | **Critical Observation:** There is **no single centralized vocabulary store**. The system is **disjoint by design** — each layer serves a different pipeline stage and has its own schema, persistence, and update mechanism. --- ## 2. Layer-by-Layer Specification --- ### 2.1 Layer A — Hard-coded Keyword Lists (CryptoSentimentCalibrator) **File:** `src/sentiment_engine/nlp/sentiment_emotion.py` **Class:** `CryptoSentimentCalibrator` (lines ~200–1400) **Purpose:** Override FinBERT's traditional-finance semantics with crypto-native semantics via keyword matching. #### 2.1.1 Data Structures ```python # Four class-level constants — all list[str] CRYPTO_BULLISH_KEYWORDS: List[str] # ~1,200+ entries CRYPTO_BEARISH_KEYWORDS: List[str] # ~1,200+ entries WHALE_BULLISH_PHRASES: List[str] # ~80 entries WHALE_BEARISH_PHRASES: List[str] # ~120 entries ``` #### 2.1.2 Entry Format | Field | Description | Example | |-------|-------------|---------| | **Keyword** | Single token or compound phrase with `.` as space placeholder | `"golden.cross"`, `"whale.accumulation"`, `"surge"` | | **Compound phrases** | Also duplicated as space-separated strings at list end | `"golden cross"`, `"whale accumulation"`, `"all time high"` | **Note:** The `.` separator is a convention for internal matching; at runtime, both `re.search(r'\b' + re.escape(kw) + r'\b', text_lower)` (for single tokens) and simple `phrase in text_lower` (for whale phrases) are used. #### 2.1.3 Categories Covered (Bullish) | Category | Example Keywords | |----------|-----------------| | Price action | `surge`, `pump`, `moon`, `rally`, `breakout`, `ath`, `higher.high` | | Inflows/accumulation | `outflow`, `whale.withdrawal`, `cold.storage`, `accumulation`, `hodl` | | Institutional/ETF | `etf`, `spot.etf`, `blackrock`, `fidelity`, `microstrategy`, `institutional.adoption` | | Exchange/listing | `listing`, `tier1.listing`, `binance.listing`, `coinbase.listing` | | Partnerships/dev | `partnership`, `integration`, `ecosystem.growth`, `developer.activity`, `grant` | | Technical indicators | `golden.cross`, `macd.crossover`, `rsi.oversold`, `support.held`, `200.day` | | On-chain | `whale.accumulation`, `exchange.outflow`, `balance.decreasing`, `staking`, `hashrate.up` | | DeFi/yield | `yield`, `apy`, `tvl.growth`, `protocol.revenue`, `buyback`, `token.burn` | | Macro/narrative | `halving`, `supply.shock`, `inflation.hedge`, `rate.cut`, `fed.pivot`, `risk.on` | | Sentiment/social | `fomo`, `euphoria`, `optimism`, `greed`, `social.dominance`, `trending` | #### 2.1.4 Categories Covered (Bearish) | Category | Example Keywords | |----------|-----------------| | Price action | `crash`, `dump`, `capitulation`, `panic`, `bear.market`, `lower.high`, `free.fall` | | Liquidations | `liquidation`, `cascade.liquidation`, `long.liquidation`, `margin.call`, `rekt` | | Hacks/security | `hack`, `exploit`, `rug`, `rugpull`, `stolen`, `vulnerability`, `flash.loan.attack` | | Depeg/stablecoin | `depeg`, `stablecoin.depeg`, `peg.broken`, `reserve.shortfall`, `undercollateralized` | | Outflows/selling | `inflow`, `exchange.inflow`, `balance.increasing`, `whale.deposit`, `profit.taking`, `paper.hands` | | Regulatory | `ban`, `lawsuit`, `sec.enforcement`, `crackdown`, `delist`, `wells.notice`, `cease.and.desist` | | Bankruptcy | `bankruptcy`, `insolvency`, `bank.run`, `withdrawal.spike`, `ftx`, `celcius`, `terra` | | Technical | `death.cross`, `macd.bearish`, `rsi.overbought`, `resistance.held`, `head.and.shoulders` | | On-chain bearish | `whale.selling`, `exchange.inflow`, `unstaking`, `hashrate.down`, `miner.capitulation` | | DeFi issues | `tvl.drop`, `protocol.exploit`, `bad.debt`, `unlock`, `token.unlock`, `dilution` | | Macro risk-off | `rate.hike`, `fed.hawkish`, `tightening`, `recession`, `inflation.high`, `dxy.up`, `risk.off` | | Sentiment/social | `fud`, `fear`, `capitulation`, `despair`, `anger`, `narrative.broken`, `thesis.invalidated` | #### 2.1.5 Whale Action Phrases (Context-Dependent) | List | Weight | Example Phrases | |------|--------|-----------------| | `WHALE_BULLISH_PHRASES` | 5× | `"whale buys"`, `"whale accumulates"`, `"whale loads"`, `"smart.money.accumulating"`, `"whale.absorbing"` | | `WHALE_BEARISH_PHRASES` | 5× | `"whale sells"`, `"whale dumps"`, `"whale distributes"`, `"whale takes profit"`, `"smart.money.selling"`, `"profit taking"` | **Weighting:** Whale phrases contribute `count * 5` to the directional score vs. `count * 1` for standard keywords. #### 2.1.6 Scoring Algorithm (`_get_crypto_signal`) ```python def _get_crypto_signal(text: str) -> str: text_lower = text.lower() # Whale phrases: simple substring match (higher priority) whale_bullish = sum(1 for phrase in WHALE_BULLISH_PHRASES if phrase in text_lower) whale_bearish = sum(1 for phrase in WHALE_BEARISH_PHRASES if phrase in text_lower) # Standard keywords: word-boundary regex match bullish_score = sum(1 for kw in CRYPTO_BULLISH_KEYWORDS if re.search(r'\b' + re.escape(kw) + r'\b', text_lower)) bearish_score = sum(1 for kw in CRYPTO_BEARISH_KEYWORDS if re.search(r'\b' + re.escape(kw) + r'\b', text_lower)) total_bullish = bullish_score + whale_bullish * 5 total_bearish = bearish_score + whale_bearish * 5 if total_bullish > total_bearish: return "bullish" elif total_bearish > total_bullish: return "bearish" return "neutral" ``` #### 2.1.7 Calibration Logic (`calibrate`) The calibrator **aggressively flips** FinBERT probabilities when crypto keywords disagree: | Crypto Signal | FinBERT Signal | Action | |---------------|----------------|--------| | bullish | bearish | Force `[0.05, neu, 0.95-neu]` | | bearish | bullish | Force `[0.95, neu, 0.05]` | | bullish | neutral | Force strong bullish | | bearish | neutral | Force strong bearish | | neutral | *any* | Force neutral (average pos/neg) | | bullish | bullish | Amplify bullish (+25% of diff) | | bearish | bearish | Amplify bearish (+50% of diff) | | *any* | weak (|diff|<0.4) | Trust crypto signal, swap pos/neg | **Key invariant:** Crypto keyword signal **always wins** when FinBERT is uncertain (|pos-neg| < 0.4). #### 2.1.8 Update Mechanism - **Add/modify:** Edit Python source → rebuild container → redeploy - **No hot-reload:** Lists are class constants loaded at import time - **Version control:** Git history tracks all changes - **Testing:** `vocab_test_cases.json` provides 200+ regression test cases --- ### 2.2 Layer B — Asset Alias Maps **File:** `config/asset_aliases.yaml` **Loaded by:** `AssetMapper.__init__()` → `EntityExtractor` **Purpose:** Map free-text mentions (names, symbols, people) → canonical ticker IDs. #### 2.2.1 Schema ```yaml aliases: "ALIAS_UPPERCASE": "CANONICAL_TICKER" # e.g. "BITCOIN": "BTC" "ETHEREUM": "ETH" "VITALIK": "ETH" "CZ": "BNB" ``` #### 2.2.2 Entry Types | Alias Type | Examples | Confidence | |------------|----------|------------| | Symbol variants | `BTC`, `XBT` → `BTC` | 0.95 | | Full names | `BITCOIN`, `ETHEREUM` → `BTC`, `ETH` | 0.95 | | Person → asset | `VITALIK` → `ETH`, `SAYLOR` → `BTC`, `ELON` → `DOGE` | 0.7–0.9 | | Stablecoins | `TETHER` → `USDT`, `CIRCLE` → `USDC` | 0.95 | | Memes | `SHIBA` → `SHIB`, `PEPE` → `PEPE` | 0.95 | #### 2.2.3 Resolution Logic (`AssetMapper.map_ticker`) 1. Direct alias match (uppercase) → confidence 0.95 2. Known entity exact match → confidence 0.9 3. Fuzzy match (rapidfuzz, cutoff 85) → confidence 0.8 × similarity 4. No match → return as-is, confidence 0.5 #### 2.2.4 Update Mechanism - Edit YAML → hot-reload on next `AssetMapper` instantiation (no code deploy) - Used by both rule-based extraction (`extract_aliases`) and NER post-processing --- ### 2.3 Layer C — Known Entities Registry **File:** `config/known_entities.yaml` **Loaded by:** `AssetMapper._load_known_entities()` **Purpose:** Rich metadata for canonical assets. #### 2.3.1 Schema ```yaml entities: BTC: name: "Bitcoin" type: "crypto" # crypto | stablecoin | defi | oracle | etc. chain: "bitcoin" contracts: [] # empty for native assets market_cap_rank: 1 ETH: name: "Ethereum" type: "crypto" chain: "ethereum" contracts: ["0xC02aaA39b223FE8D0A0e5C4F27eAD9083C756Cc2"] # WETH market_cap_rank: 2 ``` #### 2.3.2 Fields | Field | Type | Required | Description | |-------|------|----------|-------------| | `name` | str | Yes | Human-readable name | | `type` | str | Yes | Asset category (crypto, stablecoin, defi, oracle, etc.) | | `chain` | str | Yes | Native blockchain | | `contracts` | list[str] | No | Contract addresses (for wrapped/bridged versions) | | `market_cap_rank` | int | No | Coingecko-style rank | #### 2.3.3 Usage - **Contract resolution:** `AssetMapper.map_contract(address)` → matches against `contracts` list - **Fuzzy ticker match:** `rapidfuzz` against entity keys - **Entity enrichment:** `EntityExtraction.canonical_name` populated from `name` #### 2.3.4 Update Mechanism - Edit YAML → hot-reload on next `AssetMapper` instantiation - No code changes required --- ### 2.4 Layer D — Source Credibility Registry **File:** `config/source_credibility.yaml` **Loaded by:** `CatalogueManager._sync_from_config()` → `CredibilityScorer.load_registry()` **Purpose:** Per-source base credibility and relevance for weighting signals. #### 2.4.1 Schema ```yaml sources: - source_id: "rss:coindesk.com" name: "CoinDesk" url: "https://www.coindesk.com" source_type: "news" # news | research | exchange_ann | social | regulatory base_credibility: 0.85 # 0-1 static prior relevance: 0.9 # 0-1 crypto relevance enabled: true ``` #### 2.4.2 Fields | Field | Type | Range | Description | |-------|------|-------|-------------| | `source_id` | str | — | Unique ID (format: `{connector}:{identifier}`) | | `name` | str | — | Display name | | `url` | str | — | Base URL | | `source_type` | enum | news, research, exchange_ann, social, regulatory | Category for grouping | | `base_credibility` | float | [0,1] | Static prior (updated dynamically at runtime) | | `relevance` | float | [0,1] | Domain relevance to crypto markets | | `enabled` | bool | — | Whether to ingest from this source | #### 2.4.3 Runtime Dynamics - **Current credibility** (`current_credibility`) stored in DuckDB, updated by: - Fetch success/failure rates - Event outcome feedback (`confirmed` +0.02, `false_positive` -0.05, `missed` -0.03) - Time decay (half-life 30 days, min 0.1) - **Composite credibility** = `current_credibility` × `relevance` × source-type multiplier #### 2.4.4 Update Mechanism - YAML edits → hot-reload via `CatalogueManager` sync (runs on init + periodic) - Runtime updates persisted to DuckDB (`data/sources.duckdb`) --- ### 2.5 Layer E — BERT Centroids (Semantic Parameter Scoring) **Files:** `config/centroids/{fear_state,greed_state,hype_velocity,pub_velocity,pump_score,dump_score}.npy` **Managed by:** `CentroidManager` (`scoring/centroids.py`) **Purpose:** Provide semantic "meaning" for 6 scoring parameters via embedding similarity. #### 2.5.1 Structure | Parameter | File | Dimension | Description | |-----------|------|-----------|-------------| | `fear_state` | `fear_state.npy` | 768 (FinBERT) | Fear/panic semantic direction | | `greed_state` | `greed_state.npy` | 768 | Greed/FOMO semantic direction | | `hype_velocity` | `hype_velocity.npy` | 768 | Hype acceleration semantic direction | | `pub_velocity` | `pub_velocity.npy` | 768 | Publication velocity semantic direction | | `pump_score` | `pump_score.npy` | 768 | Pump/manipulation semantic direction | | `dump_score` | `dump_score.npy` | 768 | Dump/crash semantic direction | #### 2.5.2 Building Process (`_build_centroids`) ```python async def _build_centroids(self): # Current implementation: PLACEHOLDER (random unit vectors) for param in PARAMETERS: self._centroids[param] = np.random.randn(768).astype(np.float32) self._centroids[param] /= np.linalg.norm(self._centroids[param]) ``` **TODO (per code comments):** Build from keyword lists in `SENTIMENT_SPEC_IMPLEMENT_GUIDE.md`: 1. Collect keyword lists per parameter 2. Encode each keyword/sentence via `encoder` (e5-large-v2) 3. Average embeddings → unit vector centroid 4. Save to `.npy` #### 2.5.3 Scoring Usage (`_refine_with_centroids`) ```python embedding = self._get_text_embedding(combined_text) # e5-large-v2 similarity = centroid_manager.compute_similarity(embedding, param_name) centroid_score = (similarity + 1.0) / 2.0 # map [-1,1] → [0,1] params[param_name] = 0.7 * current_value + 0.3 * centroid_score ``` **Weight:** 30% centroid similarity, 70% signal-processor value. #### 2.5.4 Update Mechanism - **Current:** Placeholder — random vectors on first init if `.npy` missing - **Production:** Re-run `_build_centroids` with trained encoder → overwrite `.npy` files - **No hot-reload:** Centroids loaded once at `ScoringEngine.initialize()` --- ### 2.6 Layer F — Labeling Schema & Guidelines **File:** `labeling_pipeline.py` (lines 1–400+) **Purpose:** Define ground-truth label space for supervised training/annotation. #### 2.6.1 Sentiment Labels (3-class) | Label | Value | Description | |-------|-------|-------------| | `BEARISH` | 0 | Explicit negative price expectation | | `BULLISH` | 1 | Explicit positive price expectation | | `NEUTRAL` | 2 | No clear directional bias | **Guidelines (from `LABELING_GUIDELINES`):** | Label | Explicit Keywords | Technical | Fundamental | Emoji | |-------|------------------|-----------|-------------|-------| | BULLISH | "moon", "pump", "accumulate", "to $100k" | "golden cross", "breakout", "higher highs" | "institutional adoption", "ETF approval", "whale accumulation" | 🚀 📈 💎 🙌 🌙 | | BEARISH | "crash incoming", "dump it", "top is in" | "death cross", "breakdown", "lower high" | "SEC lawsuit", "exchange hack", "regulation ban" | 📉 😭 💀 🩸 🧻 | | NEUTRAL | "BTC at $50k", "market consolidating" | — | — | — | #### 2.6.2 Event Types (12-class) | Index | Label | Description | |-------|-------|-------------| | 0 | `listing` | New exchange listing, token debut | | 1 | `delisting` | Removal from exchange | | 2 | `hack` | Exploit, drain, theft, vulnerability | | 3 | `regulatory` | SEC, CFTC, lawsuits, regulation | | 4 | `governance` | DAO votes, proposals, treasury | | 5 | `upgrade` | Hard fork, mainnet, protocol upgrade | | 6 | `partnership` | Integration, collaboration, alliance | | 7 | `earnings` | Revenue, profit, financial results | | 8 | `macro` | Fed, rates, CPI, GDP, employment | | 9 | `liquidation` | Margin calls, cascade liquidations | | 10 | `whale` | Large transfers, accumulation, distribution | | 11 | `manipulation` | Wash trading, spoofing, pump & dump | #### 2.6.3 Emotion Types (6-class) | Label | Keywords | |-------|----------| | `joy` | moon, pump, breakout, profit, gains, success | | `fear` | crash, hack, panic, worry, risk | | `anger` | rug, scam, fraud, manipulation, unfair | | `greed` | fomo, ape, yolo, leverage, accumulation | | `sadness` | loss, rekt, down, bear, pain | | `neutral` | sideways, stable, consolidating, range | #### 2.6.4 Entity Types `TICKER`, `CONTRACT`, `PROTOCOL`, `EXCHANGE`, `PERSON`, `CHAIN`, `ORG` #### 2.6.5 Real Events (Ground Truth) `REAL_EVENTS` list in `labeling_pipeline.py` — 50+ manually labeled examples with `text`, `label_id`, `event_type`. #### 2.6.6 Update Mechanism - Edit Python enums/docstrings → rebuild - `REAL_EVENTS` extended manually for regression testing - Used by `LabelingPipelineRunner` for automated annotation --- ## 3. Pipeline Flow — How Vocabulary Flows Through the System ``` ┌─────────────────────────────────────────────────────────────────────────────────┐ │ SENTIMENT ENGINE VOCABULARY FLOW │ └─────────────────────────────────────────────────────────────────────────────────┘ RAW TEXT INPUT │ ▼ ┌─────────────────────────────────────────────────────────────────────────────┐ │ ENTITY EXTRACTION (EntityExtractor) │ │ • Ticker regex: \$?[A-Za-z]{2,10}\b │ │ • Contract regex: 0x[a-fA-F0-9]{40} | base58 │ │ • Alias lookup: Layer B (asset_aliases.yaml) + Layer C (known_entities) │ │ • NER (spaCy): ORG, PRODUCT, GPE, PERSON → fuzzy map to tickers │ │ Output: List[EntityExtraction{asset_id, mention_span, confidence, type}] │ └─────────────────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────────────────┐ │ SENTIMENT & EMOTION ANALYSIS (SentimentEmotionAnalyzer) │ │ • FinBERT (ONNX/PyTorch/Mock) → [neg, neu, pos] probs │ │ • CryptoSentimentCalibrator.calibrate(text, probs) ← LAYER A KEYWORDS │ │ - _get_crypto_signal() uses CRYPTO_BULLISH/BEARISH_KEYWORDS │ │ - WHALE_*_PHRASES weighted 5× │ │ - Word-boundary regex for standard, substring for whale phrases │ │ • Emotion model (DistilRoBERTa) → 6-class emotions │ │ • Heuristic fallback if models unavailable │ │ Output: SentimentScores(polarity, confidence, pos/neg/neu), EmotionScores │ └─────────────────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────────────────┐ │ EVENT CLASSIFICATION (EventClassifier) │ │ • BERT classifier → 12-class event type │ │ • Uses Layer F label schema (EVENT_LABELS) │ └─────────────────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────────────────┐ │ CREDIBILITY SCORING (CredibilityScorer) │ │ • Source base_credibility from Layer D (source_credibility.yaml) │ │ • Cross-source corroboration (in-memory cache) │ │ • Temporal decay (half-life 30 days) │ │ Output: CredibilityScore(composite, components) │ └─────────────────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────────────────┐ │ SIGNAL PROCESSING (SignalProcessor) │ │ • fear_state = f(neg_sentiment, fear_emotion, event_fear) │ │ • greed_state = f(pos_sentiment, greed_emotion, event_greed) │ │ • pump_score = f(greed, joy, pos_events, intensity) │ │ • dump_score = f(fear, anger, neg_events, intensity) │ │ • VelocityComputer → hype_velocity, pub_velocity │ │ • TemporalDecay (Layer E scoring.halflife_minutes) │ │ Output: AssetSentiment per asset │ └─────────────────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────────────────┐ │ CENTROID REFINEMENT (ScoringEngine._refine_with_centroids) ← LAYER E │ │ • Embed combined entity+event text via e5-large-v2 │ │ • Cosine similarity to 6 parameter centroids (Layer E .npy files) │ │ • Blend: 70% signal, 30% centroid │ └─────────────────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────────────────┐ │ AGGREGATION (Aggregator) │ │ • Asset → Industry (Layer C asset_industry_map.yaml) │ │ • Industry → Market │ │ • Decay at each level (asset 30m, industry 60m, market 120m half-life) │ │ Output: SentimentOutput(market, industries, assets) │ └─────────────────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────────────────┐ │ TRADING INTEGRATION │ │ • ACB signals: market_sentiment_state, fear_state, greed_state, │ │ hype_velocity, aggregate_pump_risk │ │ • Book health veto: pump_score > 75 │ │ • AlphaExitV7: dump_score > 70, fear_state > 80 │ └─────────────────────────────────────────────────────────────────────────────┘ ``` --- ## 4. Consistency & Centralization Analysis ### 4.1 Current State: DISJOINT | Aspect | Status | Detail | |--------|--------|--------| | **Single source of truth** | ❌ No | 6 independent stores with different schemas | | **Unified ID space** | ❌ No | Keywords (strings), aliases (ticker→ticker), entities (ticker→metadata), sources (source_id), centroids (param name), labels (enum values) | | **Versioning** | Partial | Git for code (Layer A, F), file mtime for YAML (B, C, D), file mtime for .npy (E) | | **Audit trail** | Partial | Git for code; DuckDB audit log for source credibility (D); none for centroids (E) | | **Hot-reload** | Mixed | YAML (B, C, D): yes; Python constants (A, F): no; .npy (E): no | | **Validation** | Minimal | `vocab_test_cases.json` tests Layer A only; no cross-layer validation | ### 4.2 Duplication & Drift Risks | Risk | Location | Example | |------|----------|---------| | **Keyword ↔ Label drift** | Layer A vs Layer F | `CRYPTO_BULLISH_KEYWORDS` contains "moon" but `LABELING_GUIDELINES` lists "moon" under BULLISH emoji — consistent now, but no enforcement | | **Alias ↔ Entity drift** | Layer B vs Layer C | `asset_aliases.yaml` has "VITALIK" → "ETH"; `known_entities.yaml` has ETH entry — if one updated without other, resolution breaks | | **Centroid ↔ Keyword drift** | Layer E vs Layer A | Centroids built from keywords (TODO) but currently random; if keywords change, centroids stale | | **Source credibility ↔ Event outcome** | Layer D vs Labeling | `false_positive` event outcome adjusts credibility but event labels from Layer F — no automated loop | --- ## 5. Recommendations for Centralization ### 5.1 Immediate (Low Effort) 1. **Single Vocabulary Registry** — Create `config/vocabulary.yaml` with: ```yaml sentiment_keywords: bullish: [...] bearish: [...] whale_bullish: [...] whale_bearish: [...] asset_aliases: {...} # merge Layer B known_entities: {...} # merge Layer C source_credibility: [...] # merge Layer D labeling_schema: # mirror Layer F sentiment: [BEARISH, BULLISH, NEUTRAL] events: [...] emotions: [...] ``` 2. **Runtime Loader** — `VocabularyRegistry` class loading YAML + `.npy` centroids, exposing typed accessors. 3. **Validation Tests** — Cross-layer consistency checks: - Every alias target exists in known_entities - Every whale phrase keyword appears in corresponding bullish/bearish list - Centroid rebuild script reads from `vocabulary.yaml` keyword lists ### 5.2 Medium Term 4. **Centroid Auto-Rebuild** — On vocabulary change, trigger centroid recomputation via encoder. 5. **Provenance Tracking** — Add `source: "keyword_list" | "centroid" | "heuristic"` to every score component. 6. **A/B Testing Framework** — Compare keyword-only vs. centroid-only vs. blended scoring. ### 5.3 Long Term 7. **Learned Vocabulary** — Replace hard-coded lists with learned token importance (attention weights, SHAP values) from fine-tuned model. 8. **Semantic Versioning** — `vocabulary.yaml` with `version: "2.1.0"`, migration scripts for schema changes. --- ## 6. File Inventory (Absolute Paths) | Layer | File | Lines | Size | Last Modified | |-------|------|-------|------|---------------| | A | `/mnt/dolphinng5_predict/sentiment_engine/src/sentiment_engine/nlp/sentiment_emotion.py` | ~1,776 | ~68 KB | 2026-07-xx | | B | `/mnt/dolphinng5_predict/sentiment_engine/config/asset_aliases.yaml` | ~60 | 1.1 KB | 2026-07-xx | | C | `/mnt/dolphinng5_predict/sentiment_engine/config/known_entities.yaml` | ~55 | 1.8 KB | 2026-07-xx | | D | `/mnt/dolphinng5_predict/sentiment_engine/config/source_credibility.yaml` | ~70 | 2.9 KB | 2026-07-xx | | E | `/mnt/dolphinng5_predict/sentiment_engine/config/centroids/*.npy` (6 files) | — | 3.1 KB each | 2026-07-xx | | F | `/mnt/dolphinng5_predict/sentiment_engine/labeling_pipeline.py` | ~1,000+ | ~48 KB | 2026-07-xx | | Config | `/mnt/dolphinng5_predict/sentiment_engine/config/settings.yaml` | ~180 | 7.8 KB | 2026-07-xx | | Test | `/mnt/dolphinng5_predict/sentiment_engine/vocab_test_cases.json` | ~2,000 | 47 KB | 2026-07-xx | --- ## 7. Keyword Counts (Layer A) | List | Count (approx) | Unique Stems | |------|----------------|--------------| | `CRYPTO_BULLISH_KEYWORDS` | 1,200+ | ~400 | | `CRYPTO_BEARISH_KEYWORDS` | 1,200+ | ~400 | | `WHALE_BULLISH_PHRASES` | 80 | 80 | | `WHALE_BEARISH_PHRASES` | 120 | 120 | | **Total** | **~2,600** | **~1,000** | *Note: High duplication in lists (many variants: "surge", "surges", "surged", "surgeing", "surgeing").* --- ## 8. Test Coverage (Layer A) **File:** `vocab_test_cases.json` — 200+ test cases **Coverage:** Basic positive/negative, whale phrases, compound phrases, edge cases **Run:** `pytest tests/test_crypto_sentiment_calibrator.py` (if exists) or manual via `labeling_pipeline.py` --- ## 9. Open Questions / TODOs 1. **Centroid building** — `_build_centroids()` currently uses random vectors. Implement keyword-driven centroid construction per `SENTIMENT_SPEC_IMPLEMENT_GUIDE.md`. 2. **Whale phrase matching** — Currently uses simple substring (`phrase in text_lower`). Should use word-boundary regex for consistency with standard keywords. 3. **Compound phrase deduplication** — Lists contain both `"golden.cross"` and `"golden cross"`. Normalize to single representation. 4. **Multi-word n-gram storage** — No explicit n-gram store beyond compound phrases in keyword lists. Consider adding n-gram frequency tracking from corpus. 5. **Language support** — Only English (`supported_languages: ["en"]`). Keyword lists are English-only. 6. **Dynamic keyword weighting** — All keywords equal weight (1). Could learn weights from labeled data. --- ## 10. Appendices ### 10.1 Full Keyword List Excerpt (Layer A) See `sentiment_emotion.py` lines 200–1400 for complete lists. ### 10.2 Centroid Rebuild Procedure (When Implemented) ```bash # 1. Update vocabulary.yaml with new keywords # 2. Run rebuild script python -m sentiment_engine.scripts.rebuild_centroids # 3. Verify .npy files updated # 4. Restart scoring engine ``` ### 10.3 Hot-Reload Procedures | Layer | Command | |-------|---------| | B, C, D | `POST /admin/reload-catalogue` (if API exposed) or restart `CatalogueManager` | | E | Restart `ScoringEngine` (no hot-reload) | | A, F | Full container rebuild + deploy | --- **End of Specification**