Files
sentiment-engine/sentiment_engine/VOCABULARY_SCORE_STORAGE_SPEC.md
Codex c32db97d57 feat(sentiment): complete pipeline overhaul with ONNX priority + LoRA retraining
- Added 30 new sources (5 RSS + 25 Telegram) for previously ZERO-coverage assets
- Fixed model loading priority: ONNX > LoRA v2 > PyTorch > Mock
- ONNX FinBERT (pre-trained on 1.2M financial docs) now PRIMARY - best for real-world text
- LoRA v2 models trained on 518 carefully labeled samples (balanced Bearish/Bullish/Neutral)
- Emotion LoRA v2 trained with weighted loss (greed/fear 2x, joy 1.5x)
- 30 new sources: STX, FET, XTZ, ENJ, ETC, TRX, ONG, DASH, LTC, ZIL, NEAR, APT, SUI, ICP
- Early stopping (patience=3) on both LoRA trainings
- Human-in-the-loop verification CLI tool created
- Disk-conscious: save_total_limit=1, adapters 6-8MB each

Pipeline now correctly classifies:
- BTC breaks 100k → +0.54 Bullish ✅
- Major hack → -0.23 Bearish ✅
- HODL → +0.91 Bullish ✅
- Rug pull → -0.30 Bearish ✅
- SEC sues → -0.30 Bearish ✅
- ETF approval → +0.32 Bullish ✅
- Whale accumulation → +0.31 Bullish ✅

Models: ONNX FinBERT (PRIORITY 1) + LoRA v2 adapters (6-8MB each)
Training data: 518 carefully labeled samples (190 real + 328 synthetic)
Early stopping (patience=3) on both FinBERT and DistilRoBERTa LoRA
Emotion LoRA v2: weighted loss (greed/fear 2x, joy 1.5x) + early stopping
2026-09-27 04:34:49 +02:00

32 KiB
Raw Blame History

Sentiment Engine — Vocabulary / N-gram / Phrase Storage & Scoring Specification

Version: 1.0
Date: 2026-07-16
Scope: Complete inventory of how terms, n-grams, phrases, and their "meaning/score/impact" are stored across the sentiment engine codebase.


1. Executive Summary

The sentiment engine stores vocabulary and scoring signals in six distinct layers, each with different persistence, mutability, and semantics:

Layer Storage Format Mutability Scope Primary Use
A. Hard-coded Keyword Lists Python class constants (list[str]) Code change + deploy Crypto-specific sentiment direction (bullish/bearish/whale) FinBERT calibration override
B. Asset Alias Maps YAML (config/asset_aliases.yaml) Config reload / hot-reload Canonical ticker resolution Entity extraction → asset_id mapping
C. Known Entities Registry YAML (config/known_entities.yaml) Config reload Asset metadata (chain, contracts, market cap) Entity enrichment, contract resolution
D. Source Credibility Registry YAML (config/source_credibility.yaml) Config reload Per-source base_credibility + relevance Credibility scoring, source weighting
E. BERT Centroids NumPy .npy (config/centroids/*.npy) Rebuild via encoder Semantic similarity for 6 scoring parameters Parameter refinement via embedding similarity
F. Labeling Guidelines / Schema Python enums + docstrings (labeling_pipeline.py) Code change 3-class sentiment, 12-class event, 6-class emotion Ground-truth label definitions for training

Critical Observation: There is no single centralized vocabulary store. The system is disjoint by design — each layer serves a different pipeline stage and has its own schema, persistence, and update mechanism.


2. Layer-by-Layer Specification


2.1 Layer A — Hard-coded Keyword Lists (CryptoSentimentCalibrator)

File: src/sentiment_engine/nlp/sentiment_emotion.py
Class: CryptoSentimentCalibrator (lines ~200–1400)
Purpose: Override FinBERT's traditional-finance semantics with crypto-native semantics via keyword matching.

2.1.1 Data Structures

# Four class-level constants — all list[str]

CRYPTO_BULLISH_KEYWORDS: List[str]      # ~1,200+ entries
CRYPTO_BEARISH_KEYWORDS: List[str]      # ~1,200+ entries
WHALE_BULLISH_PHRASES: List[str]        # ~80 entries
WHALE_BEARISH_PHRASES: List[str]        # ~120 entries

2.1.2 Entry Format

Field Description Example
Keyword Single token or compound phrase with . as space placeholder "golden.cross", "whale.accumulation", "surge"
Compound phrases Also duplicated as space-separated strings at list end "golden cross", "whale accumulation", "all time high"

Note: The . separator is a convention for internal matching; at runtime, both re.search(r'\b' + re.escape(kw) + r'\b', text_lower) (for single tokens) and simple phrase in text_lower (for whale phrases) are used.

2.1.3 Categories Covered (Bullish)

Category Example Keywords
Price action surge, pump, moon, rally, breakout, ath, higher.high
Inflows/accumulation outflow, whale.withdrawal, cold.storage, accumulation, hodl
Institutional/ETF etf, spot.etf, blackrock, fidelity, microstrategy, institutional.adoption
Exchange/listing listing, tier1.listing, binance.listing, coinbase.listing
Partnerships/dev partnership, integration, ecosystem.growth, developer.activity, grant
Technical indicators golden.cross, macd.crossover, rsi.oversold, support.held, 200.day
On-chain whale.accumulation, exchange.outflow, balance.decreasing, staking, hashrate.up
DeFi/yield yield, apy, tvl.growth, protocol.revenue, buyback, token.burn
Macro/narrative halving, supply.shock, inflation.hedge, rate.cut, fed.pivot, risk.on
Sentiment/social fomo, euphoria, optimism, greed, social.dominance, trending

2.1.4 Categories Covered (Bearish)

Category Example Keywords
Price action crash, dump, capitulation, panic, bear.market, lower.high, free.fall
Liquidations liquidation, cascade.liquidation, long.liquidation, margin.call, rekt
Hacks/security hack, exploit, rug, rugpull, stolen, vulnerability, flash.loan.attack
Depeg/stablecoin depeg, stablecoin.depeg, peg.broken, reserve.shortfall, undercollateralized
Outflows/selling inflow, exchange.inflow, balance.increasing, whale.deposit, profit.taking, paper.hands
Regulatory ban, lawsuit, sec.enforcement, crackdown, delist, wells.notice, cease.and.desist
Bankruptcy bankruptcy, insolvency, bank.run, withdrawal.spike, ftx, celcius, terra
Technical death.cross, macd.bearish, rsi.overbought, resistance.held, head.and.shoulders
On-chain bearish whale.selling, exchange.inflow, unstaking, hashrate.down, miner.capitulation
DeFi issues tvl.drop, protocol.exploit, bad.debt, unlock, token.unlock, dilution
Macro risk-off rate.hike, fed.hawkish, tightening, recession, inflation.high, dxy.up, risk.off
Sentiment/social fud, fear, capitulation, despair, anger, narrative.broken, thesis.invalidated

2.1.5 Whale Action Phrases (Context-Dependent)

List Weight Example Phrases
WHALE_BULLISH_PHRASES 5× "whale buys", "whale accumulates", "whale loads", "smart.money.accumulating", "whale.absorbing"
WHALE_BEARISH_PHRASES 5× "whale sells", "whale dumps", "whale distributes", "whale takes profit", "smart.money.selling", "profit taking"

Weighting: Whale phrases contribute count * 5 to the directional score vs. count * 1 for standard keywords.

2.1.6 Scoring Algorithm (_get_crypto_signal)

def _get_crypto_signal(text: str) -> str:
    text_lower = text.lower()
    
    # Whale phrases: simple substring match (higher priority)
    whale_bullish = sum(1 for phrase in WHALE_BULLISH_PHRASES if phrase in text_lower)
    whale_bearish = sum(1 for phrase in WHALE_BEARISH_PHRASES if phrase in text_lower)
    
    # Standard keywords: word-boundary regex match
    bullish_score = sum(1 for kw in CRYPTO_BULLISH_KEYWORDS 
                        if re.search(r'\b' + re.escape(kw) + r'\b', text_lower))
    bearish_score = sum(1 for kw in CRYPTO_BEARISH_KEYWORDS 
                        if re.search(r'\b' + re.escape(kw) + r'\b', text_lower))
    
    total_bullish = bullish_score + whale_bullish * 5
    total_bearish = bearish_score + whale_bearish * 5
    
    if total_bullish > total_bearish: return "bullish"
    elif total_bearish > total_bullish: return "bearish"
    return "neutral"

2.1.7 Calibration Logic (calibrate)

The calibrator aggressively flips FinBERT probabilities when crypto keywords disagree:

Crypto Signal FinBERT Signal Action
bullish bearish Force [0.05, neu, 0.95-neu]
bearish bullish Force [0.95, neu, 0.05]
bullish neutral Force strong bullish
bearish neutral Force strong bearish
neutral any Force neutral (average pos/neg)
bullish bullish Amplify bullish (+25% of diff)
bearish bearish Amplify bearish (+50% of diff)
any weak ( diff

Key invariant: Crypto keyword signal always wins when FinBERT is uncertain (|pos-neg| < 0.4).

2.1.8 Update Mechanism

  • Add/modify: Edit Python source → rebuild container → redeploy
  • No hot-reload: Lists are class constants loaded at import time
  • Version control: Git history tracks all changes
  • Testing: vocab_test_cases.json provides 200+ regression test cases

2.2 Layer B — Asset Alias Maps

File: config/asset_aliases.yaml
Loaded by: AssetMapper.__init__() → EntityExtractor
Purpose: Map free-text mentions (names, symbols, people) → canonical ticker IDs.

2.2.1 Schema

aliases:
  "ALIAS_UPPERCASE": "CANONICAL_TICKER"
  # e.g.
  "BITCOIN": "BTC"
  "ETHEREUM": "ETH"
  "VITALIK": "ETH"
  "CZ": "BNB"

2.2.2 Entry Types

Alias Type Examples Confidence
Symbol variants BTC, XBT → BTC 0.95
Full names BITCOIN, ETHEREUM → BTC, ETH 0.95
Person → asset VITALIK → ETH, SAYLOR → BTC, ELON → DOGE 0.7–0.9
Stablecoins TETHER → USDT, CIRCLE → USDC 0.95
Memes SHIBA → SHIB, PEPE → PEPE 0.95

2.2.3 Resolution Logic (AssetMapper.map_ticker)

  1. Direct alias match (uppercase) → confidence 0.95
  2. Known entity exact match → confidence 0.9
  3. Fuzzy match (rapidfuzz, cutoff 85) → confidence 0.8 × similarity
  4. No match → return as-is, confidence 0.5

2.2.4 Update Mechanism

  • Edit YAML → hot-reload on next AssetMapper instantiation (no code deploy)
  • Used by both rule-based extraction (extract_aliases) and NER post-processing

2.3 Layer C — Known Entities Registry

File: config/known_entities.yaml
Loaded by: AssetMapper._load_known_entities()
Purpose: Rich metadata for canonical assets.

2.3.1 Schema

entities:
  BTC:
    name: "Bitcoin"
    type: "crypto"           # crypto | stablecoin | defi | oracle | etc.
    chain: "bitcoin"
    contracts: []            # empty for native assets
    market_cap_rank: 1
  ETH:
    name: "Ethereum"
    type: "crypto"
    chain: "ethereum"
    contracts: ["0xC02aaA39b223FE8D0A0e5C4F27eAD9083C756Cc2"]  # WETH
    market_cap_rank: 2

2.3.2 Fields

Field Type Required Description
name str Yes Human-readable name
type str Yes Asset category (crypto, stablecoin, defi, oracle, etc.)
chain str Yes Native blockchain
contracts list[str] No Contract addresses (for wrapped/bridged versions)
market_cap_rank int No Coingecko-style rank

2.3.3 Usage

  • Contract resolution: AssetMapper.map_contract(address) → matches against contracts list
  • Fuzzy ticker match: rapidfuzz against entity keys
  • Entity enrichment: EntityExtraction.canonical_name populated from name

2.3.4 Update Mechanism

  • Edit YAML → hot-reload on next AssetMapper instantiation
  • No code changes required

2.4 Layer D — Source Credibility Registry

File: config/source_credibility.yaml
Loaded by: CatalogueManager._sync_from_config() → CredibilityScorer.load_registry()
Purpose: Per-source base credibility and relevance for weighting signals.

2.4.1 Schema

sources:
  - source_id: "rss:coindesk.com"
    name: "CoinDesk"
    url: "https://www.coindesk.com"
    source_type: "news"           # news | research | exchange_ann | social | regulatory
    base_credibility: 0.85        # 0-1 static prior
    relevance: 0.9                # 0-1 crypto relevance
    enabled: true

2.4.2 Fields

Field Type Range Description
source_id str — Unique ID (format: {connector}:{identifier})
name str — Display name
url str — Base URL
source_type enum news, research, exchange_ann, social, regulatory Category for grouping
base_credibility float [0,1] Static prior (updated dynamically at runtime)
relevance float [0,1] Domain relevance to crypto markets
enabled bool — Whether to ingest from this source

2.4.3 Runtime Dynamics

  • Current credibility (current_credibility) stored in DuckDB, updated by:
    • Fetch success/failure rates
    • Event outcome feedback (confirmed +0.02, false_positive -0.05, missed -0.03)
    • Time decay (half-life 30 days, min 0.1)
  • Composite credibility = current_credibility × relevance × source-type multiplier

2.4.4 Update Mechanism

  • YAML edits → hot-reload via CatalogueManager sync (runs on init + periodic)
  • Runtime updates persisted to DuckDB (data/sources.duckdb)

2.5 Layer E — BERT Centroids (Semantic Parameter Scoring)

Files: config/centroids/{fear_state,greed_state,hype_velocity,pub_velocity,pump_score,dump_score}.npy
Managed by: CentroidManager (scoring/centroids.py)
Purpose: Provide semantic "meaning" for 6 scoring parameters via embedding similarity.

2.5.1 Structure

Parameter File Dimension Description
fear_state fear_state.npy 768 (FinBERT) Fear/panic semantic direction
greed_state greed_state.npy 768 Greed/FOMO semantic direction
hype_velocity hype_velocity.npy 768 Hype acceleration semantic direction
pub_velocity pub_velocity.npy 768 Publication velocity semantic direction
pump_score pump_score.npy 768 Pump/manipulation semantic direction
dump_score dump_score.npy 768 Dump/crash semantic direction

2.5.2 Building Process (_build_centroids)

async def _build_centroids(self):
    # Current implementation: PLACEHOLDER (random unit vectors)
    for param in PARAMETERS:
        self._centroids[param] = np.random.randn(768).astype(np.float32)
        self._centroids[param] /= np.linalg.norm(self._centroids[param])

TODO (per code comments): Build from keyword lists in SENTIMENT_SPEC_IMPLEMENT_GUIDE.md:

  1. Collect keyword lists per parameter
  2. Encode each keyword/sentence via encoder (e5-large-v2)
  3. Average embeddings → unit vector centroid
  4. Save to .npy

2.5.3 Scoring Usage (_refine_with_centroids)

embedding = self._get_text_embedding(combined_text)  # e5-large-v2
similarity = centroid_manager.compute_similarity(embedding, param_name)
centroid_score = (similarity + 1.0) / 2.0  # map [-1,1] → [0,1]
params[param_name] = 0.7 * current_value + 0.3 * centroid_score

Weight: 30% centroid similarity, 70% signal-processor value.

2.5.4 Update Mechanism

  • Current: Placeholder — random vectors on first init if .npy missing
  • Production: Re-run _build_centroids with trained encoder → overwrite .npy files
  • No hot-reload: Centroids loaded once at ScoringEngine.initialize()

2.6 Layer F — Labeling Schema & Guidelines

File: labeling_pipeline.py (lines 1–400+)
Purpose: Define ground-truth label space for supervised training/annotation.

2.6.1 Sentiment Labels (3-class)

Label Value Description
BEARISH 0 Explicit negative price expectation
BULLISH 1 Explicit positive price expectation
NEUTRAL 2 No clear directional bias

Guidelines (from LABELING_GUIDELINES):

Label Explicit Keywords Technical Fundamental Emoji
BULLISH "moon", "pump", "accumulate", "to $100k" "golden cross", "breakout", "higher highs" "institutional adoption", "ETF approval", "whale accumulation" 🚀 📈 💎 🙌 🌙
BEARISH "crash incoming", "dump it", "top is in" "death cross", "breakdown", "lower high" "SEC lawsuit", "exchange hack", "regulation ban" 📉 😭 💀 🩸 🧻
NEUTRAL "BTC at $50k", "market consolidating" — — —

2.6.2 Event Types (12-class)

Index Label Description
0 listing New exchange listing, token debut
1 delisting Removal from exchange
2 hack Exploit, drain, theft, vulnerability
3 regulatory SEC, CFTC, lawsuits, regulation
4 governance DAO votes, proposals, treasury
5 upgrade Hard fork, mainnet, protocol upgrade
6 partnership Integration, collaboration, alliance
7 earnings Revenue, profit, financial results
8 macro Fed, rates, CPI, GDP, employment
9 liquidation Margin calls, cascade liquidations
10 whale Large transfers, accumulation, distribution
11 manipulation Wash trading, spoofing, pump & dump

2.6.3 Emotion Types (6-class)

Label Keywords
joy moon, pump, breakout, profit, gains, success
fear crash, hack, panic, worry, risk
anger rug, scam, fraud, manipulation, unfair
greed fomo, ape, yolo, leverage, accumulation
sadness loss, rekt, down, bear, pain
neutral sideways, stable, consolidating, range

2.6.4 Entity Types

TICKER, CONTRACT, PROTOCOL, EXCHANGE, PERSON, CHAIN, ORG

2.6.5 Real Events (Ground Truth)

REAL_EVENTS list in labeling_pipeline.py — 50+ manually labeled examples with text, label_id, event_type.

2.6.6 Update Mechanism

  • Edit Python enums/docstrings → rebuild
  • REAL_EVENTS extended manually for regression testing
  • Used by LabelingPipelineRunner for automated annotation

3. Pipeline Flow — How Vocabulary Flows Through the System

┌─────────────────────────────────────────────────────────────────────────────────┐
│                        SENTIMENT ENGINE VOCABULARY FLOW                          │
└─────────────────────────────────────────────────────────────────────────────────┘

RAW TEXT INPUT
      │
      ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ ENTITY EXTRACTION (EntityExtractor)                                          │
│   • Ticker regex: \$?[A-Za-z]{2,10}\b                                       │
│   • Contract regex: 0x[a-fA-F0-9]{40} | base58                              │
│   • Alias lookup: Layer B (asset_aliases.yaml) + Layer C (known_entities)   │
│   • NER (spaCy): ORG, PRODUCT, GPE, PERSON → fuzzy map to tickers           │
│   Output: List[EntityExtraction{asset_id, mention_span, confidence, type}]  │
└─────────────────────────────────────────────────────────────────────────────┘
      │
      ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ SENTIMENT & EMOTION ANALYSIS (SentimentEmotionAnalyzer)                      │
│   • FinBERT (ONNX/PyTorch/Mock) → [neg, neu, pos] probs                     │
│   • CryptoSentimentCalibrator.calibrate(text, probs)  ← LAYER A KEYWORDS    │
│       - _get_crypto_signal() uses CRYPTO_BULLISH/BEARISH_KEYWORDS           │
│       - WHALE_*_PHRASES weighted 5×                                         │
│       - Word-boundary regex for standard, substring for whale phrases       │
│   • Emotion model (DistilRoBERTa) → 6-class emotions                        │
│   • Heuristic fallback if models unavailable                                │
│   Output: SentimentScores(polarity, confidence, pos/neg/neu), EmotionScores │
└─────────────────────────────────────────────────────────────────────────────┘
      │
      ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ EVENT CLASSIFICATION (EventClassifier)                                       │
│   • BERT classifier → 12-class event type                                   │
│   • Uses Layer F label schema (EVENT_LABELS)                                │
└─────────────────────────────────────────────────────────────────────────────┘
      │
      ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ CREDIBILITY SCORING (CredibilityScorer)                                      │
│   • Source base_credibility from Layer D (source_credibility.yaml)          │
│   • Cross-source corroboration (in-memory cache)                            │
│   • Temporal decay (half-life 30 days)                                      │
│   Output: CredibilityScore(composite, components)                           │
└─────────────────────────────────────────────────────────────────────────────┘
      │
      ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ SIGNAL PROCESSING (SignalProcessor)                                          │
│   • fear_state = f(neg_sentiment, fear_emotion, event_fear)                 │
│   • greed_state = f(pos_sentiment, greed_emotion, event_greed)              │
│   • pump_score = f(greed, joy, pos_events, intensity)                       │
│   • dump_score = f(fear, anger, neg_events, intensity)                      │
│   • VelocityComputer → hype_velocity, pub_velocity                          │
│   • TemporalDecay (Layer E scoring.halflife_minutes)                        │
│   Output: AssetSentiment per asset                                          │
└─────────────────────────────────────────────────────────────────────────────┘
      │
      ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ CENTROID REFINEMENT (ScoringEngine._refine_with_centroids)  ← LAYER E       │
│   • Embed combined entity+event text via e5-large-v2                        │
│   • Cosine similarity to 6 parameter centroids (Layer E .npy files)         │
│   • Blend: 70% signal, 30% centroid                                         │
└─────────────────────────────────────────────────────────────────────────────┘
      │
      ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ AGGREGATION (Aggregator)                                                     │
│   • Asset → Industry (Layer C asset_industry_map.yaml)                      │
│   • Industry → Market                                                       │
│   • Decay at each level (asset 30m, industry 60m, market 120m half-life)   │
│   Output: SentimentOutput(market, industries, assets)                       │
└─────────────────────────────────────────────────────────────────────────────┘
      │
      ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ TRADING INTEGRATION                                                            │
│   • ACB signals: market_sentiment_state, fear_state, greed_state,           │
│     hype_velocity, aggregate_pump_risk                                      │
│   • Book health veto: pump_score > 75                                       │
│   • AlphaExitV7: dump_score > 70, fear_state > 80                           │
└─────────────────────────────────────────────────────────────────────────────┘

4. Consistency & Centralization Analysis

4.1 Current State: DISJOINT

Aspect Status Detail
Single source of truth ❌ No 6 independent stores with different schemas
Unified ID space ❌ No Keywords (strings), aliases (ticker→ticker), entities (ticker→metadata), sources (source_id), centroids (param name), labels (enum values)
Versioning Partial Git for code (Layer A, F), file mtime for YAML (B, C, D), file mtime for .npy (E)
Audit trail Partial Git for code; DuckDB audit log for source credibility (D); none for centroids (E)
Hot-reload Mixed YAML (B, C, D): yes; Python constants (A, F): no; .npy (E): no
Validation Minimal vocab_test_cases.json tests Layer A only; no cross-layer validation

4.2 Duplication & Drift Risks

Risk Location Example
Keyword ↔ Label drift Layer A vs Layer F CRYPTO_BULLISH_KEYWORDS contains "moon" but LABELING_GUIDELINES lists "moon" under BULLISH emoji — consistent now, but no enforcement
Alias ↔ Entity drift Layer B vs Layer C asset_aliases.yaml has "VITALIK" → "ETH"; known_entities.yaml has ETH entry — if one updated without other, resolution breaks
Centroid ↔ Keyword drift Layer E vs Layer A Centroids built from keywords (TODO) but currently random; if keywords change, centroids stale
Source credibility ↔ Event outcome Layer D vs Labeling false_positive event outcome adjusts credibility but event labels from Layer F — no automated loop

5. Recommendations for Centralization

5.1 Immediate (Low Effort)

  1. Single Vocabulary Registry — Create config/vocabulary.yaml with:

    sentiment_keywords:
      bullish: [...]
      bearish: [...]
      whale_bullish: [...]
      whale_bearish: [...]
    asset_aliases: {...}           # merge Layer B
    known_entities: {...}          # merge Layer C
    source_credibility: [...]      # merge Layer D
    labeling_schema:               # mirror Layer F
      sentiment: [BEARISH, BULLISH, NEUTRAL]
      events: [...]
      emotions: [...]
    
  2. Runtime Loader — VocabularyRegistry class loading YAML + .npy centroids, exposing typed accessors.

  3. Validation Tests — Cross-layer consistency checks:

    • Every alias target exists in known_entities
    • Every whale phrase keyword appears in corresponding bullish/bearish list
    • Centroid rebuild script reads from vocabulary.yaml keyword lists

5.2 Medium Term

  1. Centroid Auto-Rebuild — On vocabulary change, trigger centroid recomputation via encoder.

  2. Provenance Tracking — Add source: "keyword_list" | "centroid" | "heuristic" to every score component.

  3. A/B Testing Framework — Compare keyword-only vs. centroid-only vs. blended scoring.

5.3 Long Term

  1. Learned Vocabulary — Replace hard-coded lists with learned token importance (attention weights, SHAP values) from fine-tuned model.

  2. Semantic Versioning — vocabulary.yaml with version: "2.1.0", migration scripts for schema changes.


6. File Inventory (Absolute Paths)

Layer File Lines Size Last Modified
A /mnt/dolphinng5_predict/sentiment_engine/src/sentiment_engine/nlp/sentiment_emotion.py ~1,776 ~68 KB 2026-07-xx
B /mnt/dolphinng5_predict/sentiment_engine/config/asset_aliases.yaml ~60 1.1 KB 2026-07-xx
C /mnt/dolphinng5_predict/sentiment_engine/config/known_entities.yaml ~55 1.8 KB 2026-07-xx
D /mnt/dolphinng5_predict/sentiment_engine/config/source_credibility.yaml ~70 2.9 KB 2026-07-xx
E /mnt/dolphinng5_predict/sentiment_engine/config/centroids/*.npy (6 files) — 3.1 KB each 2026-07-xx
F /mnt/dolphinng5_predict/sentiment_engine/labeling_pipeline.py ~1,000+ ~48 KB 2026-07-xx
Config /mnt/dolphinng5_predict/sentiment_engine/config/settings.yaml ~180 7.8 KB 2026-07-xx
Test /mnt/dolphinng5_predict/sentiment_engine/vocab_test_cases.json ~2,000 47 KB 2026-07-xx

7. Keyword Counts (Layer A)

List Count (approx) Unique Stems
CRYPTO_BULLISH_KEYWORDS 1,200+ ~400
CRYPTO_BEARISH_KEYWORDS 1,200+ ~400
WHALE_BULLISH_PHRASES 80 80
WHALE_BEARISH_PHRASES 120 120
Total ~2,600 ~1,000

Note: High duplication in lists (many variants: "surge", "surges", "surged", "surgeing", "surgeing").


8. Test Coverage (Layer A)

File: vocab_test_cases.json — 200+ test cases
Coverage: Basic positive/negative, whale phrases, compound phrases, edge cases
Run: pytest tests/test_crypto_sentiment_calibrator.py (if exists) or manual via labeling_pipeline.py


9. Open Questions / TODOs

  1. Centroid building — _build_centroids() currently uses random vectors. Implement keyword-driven centroid construction per SENTIMENT_SPEC_IMPLEMENT_GUIDE.md.

  2. Whale phrase matching — Currently uses simple substring (phrase in text_lower). Should use word-boundary regex for consistency with standard keywords.

  3. Compound phrase deduplication — Lists contain both "golden.cross" and "golden cross". Normalize to single representation.

  4. Multi-word n-gram storage — No explicit n-gram store beyond compound phrases in keyword lists. Consider adding n-gram frequency tracking from corpus.

  5. Language support — Only English (supported_languages: ["en"]). Keyword lists are English-only.

  6. Dynamic keyword weighting — All keywords equal weight (1). Could learn weights from labeled data.


10. Appendices

10.1 Full Keyword List Excerpt (Layer A)

See sentiment_emotion.py lines 200–1400 for complete lists.

10.2 Centroid Rebuild Procedure (When Implemented)

# 1. Update vocabulary.yaml with new keywords
# 2. Run rebuild script
python -m sentiment_engine.scripts.rebuild_centroids
# 3. Verify .npy files updated
# 4. Restart scoring engine

10.3 Hot-Reload Procedures

Layer Command
B, C, D POST /admin/reload-catalogue (if API exposed) or restart CatalogueManager
E Restart ScoringEngine (no hot-reload)
A, F Full container rebuild + deploy

End of Specification