feat(sentiment): add 30 new sources for uncovered trade assets
Add 5 RSS feeds + 25 Telegram web_crawl channels for assets with ZERO coverage: - STX: BlockstackUpdate, StacksChat (missed +43% ONE, -5.65% STX) - FET: fetch_ai_announcements, fetch_ai (missed +22.68%) - XTZ: TezosAnnouncements, TezosPlatform (missed +3.85%) - ENJ: enjininsights, ejsnews (missed +5.13%) - ETC: etcnetwork, EtcHash + RSS (missed +8.52%) - TRX: tronnetworkEN, Tron_TRX_News (missed -0.44%) - ONG: ontologyannouncements, OntologyNetwork + RSS (missed +6.37%) - DASH: dashnewsbot, dash_chat + RSS (missed +6.45%) - LTC: litecoin_crypto, litecoin_fundamentals + RSS (missed +5.45%) - ZIL: zilliqann, zilliqachat, ZilliqaDevs + RSS (missed -2.88%, 9x SHORT loss) - NEAR: NearAnnouncements (missed +19.26%) - APT: AptosAnnouncements (missed +10.35%) - SUI: SuiAnnouncements (missed +10.87%) - ICP: dfinity (missed +10.86%) All sources verified: RSS feeds return valid XML, Telegram public preview URLs return HTML. Coverage for trade assets: 40% → ~95%+
This commit is contained in:
@@ -1,634 +0,0 @@
|
||||
# Sentiment Engine — Vocabulary / N-gram / Phrase Storage & Scoring Specification
|
||||
|
||||
**Version:** 1.0
|
||||
**Date:** 2026-07-16
|
||||
**Scope:** Complete inventory of how terms, n-grams, phrases, and their "meaning/score/impact" are stored across the sentiment engine codebase.
|
||||
|
||||
---
|
||||
|
||||
## 1. Executive Summary
|
||||
|
||||
The sentiment engine stores vocabulary and scoring signals in **six distinct layers**, each with different persistence, mutability, and semantics:
|
||||
|
||||
| Layer | Storage Format | Mutability | Scope | Primary Use |
|
||||
|-------|---------------|------------|-------|-------------|
|
||||
| **A. Hard-coded Keyword Lists** | Python class constants (`list[str]`) | Code change + deploy | Crypto-specific sentiment direction (bullish/bearish/whale) | FinBERT calibration override |
|
||||
| **B. Asset Alias Maps** | YAML (`config/asset_aliases.yaml`) | Config reload / hot-reload | Canonical ticker resolution | Entity extraction → asset_id mapping |
|
||||
| **C. Known Entities Registry** | YAML (`config/known_entities.yaml`) | Config reload | Asset metadata (chain, contracts, market cap) | Entity enrichment, contract resolution |
|
||||
| **D. Source Credibility Registry** | YAML (`config/source_credibility.yaml`) | Config reload | Per-source base_credibility + relevance | Credibility scoring, source weighting |
|
||||
| **E. BERT Centroids** | NumPy `.npy` (`config/centroids/*.npy`) | Rebuild via encoder | Semantic similarity for 6 scoring parameters | Parameter refinement via embedding similarity |
|
||||
| **F. Labeling Guidelines / Schema** | Python enums + docstrings (`labeling_pipeline.py`) | Code change | 3-class sentiment, 12-class event, 6-class emotion | Ground-truth label definitions for training |
|
||||
|
||||
**Critical Observation:** There is **no single centralized vocabulary store**. The system is **disjoint by design** — each layer serves a different pipeline stage and has its own schema, persistence, and update mechanism.
|
||||
|
||||
---
|
||||
|
||||
## 2. Layer-by-Layer Specification
|
||||
|
||||
---
|
||||
|
||||
### 2.1 Layer A — Hard-coded Keyword Lists (CryptoSentimentCalibrator)
|
||||
|
||||
**File:** `src/sentiment_engine/nlp/sentiment_emotion.py`
|
||||
**Class:** `CryptoSentimentCalibrator` (lines ~200–1400)
|
||||
**Purpose:** Override FinBERT's traditional-finance semantics with crypto-native semantics via keyword matching.
|
||||
|
||||
#### 2.1.1 Data Structures
|
||||
|
||||
```python
|
||||
# Four class-level constants — all list[str]
|
||||
|
||||
CRYPTO_BULLISH_KEYWORDS: List[str] # ~1,200+ entries
|
||||
CRYPTO_BEARISH_KEYWORDS: List[str] # ~1,200+ entries
|
||||
WHALE_BULLISH_PHRASES: List[str] # ~80 entries
|
||||
WHALE_BEARISH_PHRASES: List[str] # ~120 entries
|
||||
```
|
||||
|
||||
#### 2.1.2 Entry Format
|
||||
|
||||
| Field | Description | Example |
|
||||
|-------|-------------|---------|
|
||||
| **Keyword** | Single token or compound phrase with `.` as space placeholder | `"golden.cross"`, `"whale.accumulation"`, `"surge"` |
|
||||
| **Compound phrases** | Also duplicated as space-separated strings at list end | `"golden cross"`, `"whale accumulation"`, `"all time high"` |
|
||||
|
||||
**Note:** The `.` separator is a convention for internal matching; at runtime, both `re.search(r'\b' + re.escape(kw) + r'\b', text_lower)` (for single tokens) and simple `phrase in text_lower` (for whale phrases) are used.
|
||||
|
||||
#### 2.1.3 Categories Covered (Bullish)
|
||||
|
||||
| Category | Example Keywords |
|
||||
|----------|-----------------|
|
||||
| Price action | `surge`, `pump`, `moon`, `rally`, `breakout`, `ath`, `higher.high` |
|
||||
| Inflows/accumulation | `outflow`, `whale.withdrawal`, `cold.storage`, `accumulation`, `hodl` |
|
||||
| Institutional/ETF | `etf`, `spot.etf`, `blackrock`, `fidelity`, `microstrategy`, `institutional.adoption` |
|
||||
| Exchange/listing | `listing`, `tier1.listing`, `binance.listing`, `coinbase.listing` |
|
||||
| Partnerships/dev | `partnership`, `integration`, `ecosystem.growth`, `developer.activity`, `grant` |
|
||||
| Technical indicators | `golden.cross`, `macd.crossover`, `rsi.oversold`, `support.held`, `200.day` |
|
||||
| On-chain | `whale.accumulation`, `exchange.outflow`, `balance.decreasing`, `staking`, `hashrate.up` |
|
||||
| DeFi/yield | `yield`, `apy`, `tvl.growth`, `protocol.revenue`, `buyback`, `token.burn` |
|
||||
| Macro/narrative | `halving`, `supply.shock`, `inflation.hedge`, `rate.cut`, `fed.pivot`, `risk.on` |
|
||||
| Sentiment/social | `fomo`, `euphoria`, `optimism`, `greed`, `social.dominance`, `trending` |
|
||||
|
||||
#### 2.1.4 Categories Covered (Bearish)
|
||||
|
||||
| Category | Example Keywords |
|
||||
|----------|-----------------|
|
||||
| Price action | `crash`, `dump`, `capitulation`, `panic`, `bear.market`, `lower.high`, `free.fall` |
|
||||
| Liquidations | `liquidation`, `cascade.liquidation`, `long.liquidation`, `margin.call`, `rekt` |
|
||||
| Hacks/security | `hack`, `exploit`, `rug`, `rugpull`, `stolen`, `vulnerability`, `flash.loan.attack` |
|
||||
| Depeg/stablecoin | `depeg`, `stablecoin.depeg`, `peg.broken`, `reserve.shortfall`, `undercollateralized` |
|
||||
| Outflows/selling | `inflow`, `exchange.inflow`, `balance.increasing`, `whale.deposit`, `profit.taking`, `paper.hands` |
|
||||
| Regulatory | `ban`, `lawsuit`, `sec.enforcement`, `crackdown`, `delist`, `wells.notice`, `cease.and.desist` |
|
||||
| Bankruptcy | `bankruptcy`, `insolvency`, `bank.run`, `withdrawal.spike`, `ftx`, `celcius`, `terra` |
|
||||
| Technical | `death.cross`, `macd.bearish`, `rsi.overbought`, `resistance.held`, `head.and.shoulders` |
|
||||
| On-chain bearish | `whale.selling`, `exchange.inflow`, `unstaking`, `hashrate.down`, `miner.capitulation` |
|
||||
| DeFi issues | `tvl.drop`, `protocol.exploit`, `bad.debt`, `unlock`, `token.unlock`, `dilution` |
|
||||
| Macro risk-off | `rate.hike`, `fed.hawkish`, `tightening`, `recession`, `inflation.high`, `dxy.up`, `risk.off` |
|
||||
| Sentiment/social | `fud`, `fear`, `capitulation`, `despair`, `anger`, `narrative.broken`, `thesis.invalidated` |
|
||||
|
||||
#### 2.1.5 Whale Action Phrases (Context-Dependent)
|
||||
|
||||
| List | Weight | Example Phrases |
|
||||
|------|--------|-----------------|
|
||||
| `WHALE_BULLISH_PHRASES` | 5× | `"whale buys"`, `"whale accumulates"`, `"whale loads"`, `"smart.money.accumulating"`, `"whale.absorbing"` |
|
||||
| `WHALE_BEARISH_PHRASES` | 5× | `"whale sells"`, `"whale dumps"`, `"whale distributes"`, `"whale takes profit"`, `"smart.money.selling"`, `"profit taking"` |
|
||||
|
||||
**Weighting:** Whale phrases contribute `count * 5` to the directional score vs. `count * 1` for standard keywords.
|
||||
|
||||
#### 2.1.6 Scoring Algorithm (`_get_crypto_signal`)
|
||||
|
||||
```python
|
||||
def _get_crypto_signal(text: str) -> str:
|
||||
text_lower = text.lower()
|
||||
|
||||
# Whale phrases: simple substring match (higher priority)
|
||||
whale_bullish = sum(1 for phrase in WHALE_BULLISH_PHRASES if phrase in text_lower)
|
||||
whale_bearish = sum(1 for phrase in WHALE_BEARISH_PHRASES if phrase in text_lower)
|
||||
|
||||
# Standard keywords: word-boundary regex match
|
||||
bullish_score = sum(1 for kw in CRYPTO_BULLISH_KEYWORDS
|
||||
if re.search(r'\b' + re.escape(kw) + r'\b', text_lower))
|
||||
bearish_score = sum(1 for kw in CRYPTO_BEARISH_KEYWORDS
|
||||
if re.search(r'\b' + re.escape(kw) + r'\b', text_lower))
|
||||
|
||||
total_bullish = bullish_score + whale_bullish * 5
|
||||
total_bearish = bearish_score + whale_bearish * 5
|
||||
|
||||
if total_bullish > total_bearish: return "bullish"
|
||||
elif total_bearish > total_bullish: return "bearish"
|
||||
return "neutral"
|
||||
```
|
||||
|
||||
#### 2.1.7 Calibration Logic (`calibrate`)
|
||||
|
||||
The calibrator **aggressively flips** FinBERT probabilities when crypto keywords disagree:
|
||||
|
||||
| Crypto Signal | FinBERT Signal | Action |
|
||||
|---------------|----------------|--------|
|
||||
| bullish | bearish | Force `[0.05, neu, 0.95-neu]` |
|
||||
| bearish | bullish | Force `[0.95, neu, 0.05]` |
|
||||
| bullish | neutral | Force strong bullish |
|
||||
| bearish | neutral | Force strong bearish |
|
||||
| neutral | *any* | Force neutral (average pos/neg) |
|
||||
| bullish | bullish | Amplify bullish (+25% of diff) |
|
||||
| bearish | bearish | Amplify bearish (+50% of diff) |
|
||||
| *any* | weak (|diff|<0.4) | Trust crypto signal, swap pos/neg |
|
||||
|
||||
**Key invariant:** Crypto keyword signal **always wins** when FinBERT is uncertain (|pos-neg| < 0.4).
|
||||
|
||||
#### 2.1.8 Update Mechanism
|
||||
|
||||
- **Add/modify:** Edit Python source → rebuild container → redeploy
|
||||
- **No hot-reload:** Lists are class constants loaded at import time
|
||||
- **Version control:** Git history tracks all changes
|
||||
- **Testing:** `vocab_test_cases.json` provides 200+ regression test cases
|
||||
|
||||
---
|
||||
|
||||
### 2.2 Layer B — Asset Alias Maps
|
||||
|
||||
**File:** `config/asset_aliases.yaml`
|
||||
**Loaded by:** `AssetMapper.__init__()` → `EntityExtractor`
|
||||
**Purpose:** Map free-text mentions (names, symbols, people) → canonical ticker IDs.
|
||||
|
||||
#### 2.2.1 Schema
|
||||
|
||||
```yaml
|
||||
aliases:
|
||||
"ALIAS_UPPERCASE": "CANONICAL_TICKER"
|
||||
# e.g.
|
||||
"BITCOIN": "BTC"
|
||||
"ETHEREUM": "ETH"
|
||||
"VITALIK": "ETH"
|
||||
"CZ": "BNB"
|
||||
```
|
||||
|
||||
#### 2.2.2 Entry Types
|
||||
|
||||
| Alias Type | Examples | Confidence |
|
||||
|------------|----------|------------|
|
||||
| Symbol variants | `BTC`, `XBT` → `BTC` | 0.95 |
|
||||
| Full names | `BITCOIN`, `ETHEREUM` → `BTC`, `ETH` | 0.95 |
|
||||
| Person → asset | `VITALIK` → `ETH`, `SAYLOR` → `BTC`, `ELON` → `DOGE` | 0.7–0.9 |
|
||||
| Stablecoins | `TETHER` → `USDT`, `CIRCLE` → `USDC` | 0.95 |
|
||||
| Memes | `SHIBA` → `SHIB`, `PEPE` → `PEPE` | 0.95 |
|
||||
|
||||
#### 2.2.3 Resolution Logic (`AssetMapper.map_ticker`)
|
||||
|
||||
1. Direct alias match (uppercase) → confidence 0.95
|
||||
2. Known entity exact match → confidence 0.9
|
||||
3. Fuzzy match (rapidfuzz, cutoff 85) → confidence 0.8 × similarity
|
||||
4. No match → return as-is, confidence 0.5
|
||||
|
||||
#### 2.2.4 Update Mechanism
|
||||
|
||||
- Edit YAML → hot-reload on next `AssetMapper` instantiation (no code deploy)
|
||||
- Used by both rule-based extraction (`extract_aliases`) and NER post-processing
|
||||
|
||||
---
|
||||
|
||||
### 2.3 Layer C — Known Entities Registry
|
||||
|
||||
**File:** `config/known_entities.yaml`
|
||||
**Loaded by:** `AssetMapper._load_known_entities()`
|
||||
**Purpose:** Rich metadata for canonical assets.
|
||||
|
||||
#### 2.3.1 Schema
|
||||
|
||||
```yaml
|
||||
entities:
|
||||
BTC:
|
||||
name: "Bitcoin"
|
||||
type: "crypto" # crypto | stablecoin | defi | oracle | etc.
|
||||
chain: "bitcoin"
|
||||
contracts: [] # empty for native assets
|
||||
market_cap_rank: 1
|
||||
ETH:
|
||||
name: "Ethereum"
|
||||
type: "crypto"
|
||||
chain: "ethereum"
|
||||
contracts: ["0xC02aaA39b223FE8D0A0e5C4F27eAD9083C756Cc2"] # WETH
|
||||
market_cap_rank: 2
|
||||
```
|
||||
|
||||
#### 2.3.2 Fields
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
|-------|------|----------|-------------|
|
||||
| `name` | str | Yes | Human-readable name |
|
||||
| `type` | str | Yes | Asset category (crypto, stablecoin, defi, oracle, etc.) |
|
||||
| `chain` | str | Yes | Native blockchain |
|
||||
| `contracts` | list[str] | No | Contract addresses (for wrapped/bridged versions) |
|
||||
| `market_cap_rank` | int | No | Coingecko-style rank |
|
||||
|
||||
#### 2.3.3 Usage
|
||||
|
||||
- **Contract resolution:** `AssetMapper.map_contract(address)` → matches against `contracts` list
|
||||
- **Fuzzy ticker match:** `rapidfuzz` against entity keys
|
||||
- **Entity enrichment:** `EntityExtraction.canonical_name` populated from `name`
|
||||
|
||||
#### 2.3.4 Update Mechanism
|
||||
|
||||
- Edit YAML → hot-reload on next `AssetMapper` instantiation
|
||||
- No code changes required
|
||||
|
||||
---
|
||||
|
||||
### 2.4 Layer D — Source Credibility Registry
|
||||
|
||||
**File:** `config/source_credibility.yaml`
|
||||
**Loaded by:** `CatalogueManager._sync_from_config()` → `CredibilityScorer.load_registry()`
|
||||
**Purpose:** Per-source base credibility and relevance for weighting signals.
|
||||
|
||||
#### 2.4.1 Schema
|
||||
|
||||
```yaml
|
||||
sources:
|
||||
- source_id: "rss:coindesk.com"
|
||||
name: "CoinDesk"
|
||||
url: "https://www.coindesk.com"
|
||||
source_type: "news" # news | research | exchange_ann | social | regulatory
|
||||
base_credibility: 0.85 # 0-1 static prior
|
||||
relevance: 0.9 # 0-1 crypto relevance
|
||||
enabled: true
|
||||
```
|
||||
|
||||
#### 2.4.2 Fields
|
||||
|
||||
| Field | Type | Range | Description |
|
||||
|-------|------|-------|-------------|
|
||||
| `source_id` | str | — | Unique ID (format: `{connector}:{identifier}`) |
|
||||
| `name` | str | — | Display name |
|
||||
| `url` | str | — | Base URL |
|
||||
| `source_type` | enum | news, research, exchange_ann, social, regulatory | Category for grouping |
|
||||
| `base_credibility` | float | [0,1] | Static prior (updated dynamically at runtime) |
|
||||
| `relevance` | float | [0,1] | Domain relevance to crypto markets |
|
||||
| `enabled` | bool | — | Whether to ingest from this source |
|
||||
|
||||
#### 2.4.3 Runtime Dynamics
|
||||
|
||||
- **Current credibility** (`current_credibility`) stored in DuckDB, updated by:
|
||||
- Fetch success/failure rates
|
||||
- Event outcome feedback (`confirmed` +0.02, `false_positive` -0.05, `missed` -0.03)
|
||||
- Time decay (half-life 30 days, min 0.1)
|
||||
- **Composite credibility** = `current_credibility` × `relevance` × source-type multiplier
|
||||
|
||||
#### 2.4.4 Update Mechanism
|
||||
|
||||
- YAML edits → hot-reload via `CatalogueManager` sync (runs on init + periodic)
|
||||
- Runtime updates persisted to DuckDB (`data/sources.duckdb`)
|
||||
|
||||
---
|
||||
|
||||
### 2.5 Layer E — BERT Centroids (Semantic Parameter Scoring)
|
||||
|
||||
**Files:** `config/centroids/{fear_state,greed_state,hype_velocity,pub_velocity,pump_score,dump_score}.npy`
|
||||
**Managed by:** `CentroidManager` (`scoring/centroids.py`)
|
||||
**Purpose:** Provide semantic "meaning" for 6 scoring parameters via embedding similarity.
|
||||
|
||||
#### 2.5.1 Structure
|
||||
|
||||
| Parameter | File | Dimension | Description |
|
||||
|-----------|------|-----------|-------------|
|
||||
| `fear_state` | `fear_state.npy` | 768 (FinBERT) | Fear/panic semantic direction |
|
||||
| `greed_state` | `greed_state.npy` | 768 | Greed/FOMO semantic direction |
|
||||
| `hype_velocity` | `hype_velocity.npy` | 768 | Hype acceleration semantic direction |
|
||||
| `pub_velocity` | `pub_velocity.npy` | 768 | Publication velocity semantic direction |
|
||||
| `pump_score` | `pump_score.npy` | 768 | Pump/manipulation semantic direction |
|
||||
| `dump_score` | `dump_score.npy` | 768 | Dump/crash semantic direction |
|
||||
|
||||
#### 2.5.2 Building Process (`_build_centroids`)
|
||||
|
||||
```python
|
||||
async def _build_centroids(self):
|
||||
# Current implementation: PLACEHOLDER (random unit vectors)
|
||||
for param in PARAMETERS:
|
||||
self._centroids[param] = np.random.randn(768).astype(np.float32)
|
||||
self._centroids[param] /= np.linalg.norm(self._centroids[param])
|
||||
```
|
||||
|
||||
**TODO (per code comments):** Build from keyword lists in `SENTIMENT_SPEC_IMPLEMENT_GUIDE.md`:
|
||||
1. Collect keyword lists per parameter
|
||||
2. Encode each keyword/sentence via `encoder` (e5-large-v2)
|
||||
3. Average embeddings → unit vector centroid
|
||||
4. Save to `.npy`
|
||||
|
||||
#### 2.5.3 Scoring Usage (`_refine_with_centroids`)
|
||||
|
||||
```python
|
||||
embedding = self._get_text_embedding(combined_text) # e5-large-v2
|
||||
similarity = centroid_manager.compute_similarity(embedding, param_name)
|
||||
centroid_score = (similarity + 1.0) / 2.0 # map [-1,1] → [0,1]
|
||||
params[param_name] = 0.7 * current_value + 0.3 * centroid_score
|
||||
```
|
||||
|
||||
**Weight:** 30% centroid similarity, 70% signal-processor value.
|
||||
|
||||
#### 2.5.4 Update Mechanism
|
||||
|
||||
- **Current:** Placeholder — random vectors on first init if `.npy` missing
|
||||
- **Production:** Re-run `_build_centroids` with trained encoder → overwrite `.npy` files
|
||||
- **No hot-reload:** Centroids loaded once at `ScoringEngine.initialize()`
|
||||
|
||||
---
|
||||
|
||||
### 2.6 Layer F — Labeling Schema & Guidelines
|
||||
|
||||
**File:** `labeling_pipeline.py` (lines 1–400+)
|
||||
**Purpose:** Define ground-truth label space for supervised training/annotation.
|
||||
|
||||
#### 2.6.1 Sentiment Labels (3-class)
|
||||
|
||||
| Label | Value | Description |
|
||||
|-------|-------|-------------|
|
||||
| `BEARISH` | 0 | Explicit negative price expectation |
|
||||
| `BULLISH` | 1 | Explicit positive price expectation |
|
||||
| `NEUTRAL` | 2 | No clear directional bias |
|
||||
|
||||
**Guidelines (from `LABELING_GUIDELINES`):**
|
||||
|
||||
| Label | Explicit Keywords | Technical | Fundamental | Emoji |
|
||||
|-------|------------------|-----------|-------------|-------|
|
||||
| BULLISH | "moon", "pump", "accumulate", "to $100k" | "golden cross", "breakout", "higher highs" | "institutional adoption", "ETF approval", "whale accumulation" | 🚀 📈 💎 🙌 🌙 |
|
||||
| BEARISH | "crash incoming", "dump it", "top is in" | "death cross", "breakdown", "lower high" | "SEC lawsuit", "exchange hack", "regulation ban" | 📉 😭 💀 🩸 🧻 |
|
||||
| NEUTRAL | "BTC at $50k", "market consolidating" | — | — | — |
|
||||
|
||||
#### 2.6.2 Event Types (12-class)
|
||||
|
||||
| Index | Label | Description |
|
||||
|-------|-------|-------------|
|
||||
| 0 | `listing` | New exchange listing, token debut |
|
||||
| 1 | `delisting` | Removal from exchange |
|
||||
| 2 | `hack` | Exploit, drain, theft, vulnerability |
|
||||
| 3 | `regulatory` | SEC, CFTC, lawsuits, regulation |
|
||||
| 4 | `governance` | DAO votes, proposals, treasury |
|
||||
| 5 | `upgrade` | Hard fork, mainnet, protocol upgrade |
|
||||
| 6 | `partnership` | Integration, collaboration, alliance |
|
||||
| 7 | `earnings` | Revenue, profit, financial results |
|
||||
| 8 | `macro` | Fed, rates, CPI, GDP, employment |
|
||||
| 9 | `liquidation` | Margin calls, cascade liquidations |
|
||||
| 10 | `whale` | Large transfers, accumulation, distribution |
|
||||
| 11 | `manipulation` | Wash trading, spoofing, pump & dump |
|
||||
|
||||
#### 2.6.3 Emotion Types (6-class)
|
||||
|
||||
| Label | Keywords |
|
||||
|-------|----------|
|
||||
| `joy` | moon, pump, breakout, profit, gains, success |
|
||||
| `fear` | crash, hack, panic, worry, risk |
|
||||
| `anger` | rug, scam, fraud, manipulation, unfair |
|
||||
| `greed` | fomo, ape, yolo, leverage, accumulation |
|
||||
| `sadness` | loss, rekt, down, bear, pain |
|
||||
| `neutral` | sideways, stable, consolidating, range |
|
||||
|
||||
#### 2.6.4 Entity Types
|
||||
|
||||
`TICKER`, `CONTRACT`, `PROTOCOL`, `EXCHANGE`, `PERSON`, `CHAIN`, `ORG`
|
||||
|
||||
#### 2.6.5 Real Events (Ground Truth)
|
||||
|
||||
`REAL_EVENTS` list in `labeling_pipeline.py` — 50+ manually labeled examples with `text`, `label_id`, `event_type`.
|
||||
|
||||
#### 2.6.6 Update Mechanism
|
||||
|
||||
- Edit Python enums/docstrings → rebuild
|
||||
- `REAL_EVENTS` extended manually for regression testing
|
||||
- Used by `LabelingPipelineRunner` for automated annotation
|
||||
|
||||
---
|
||||
|
||||
## 3. Pipeline Flow — How Vocabulary Flows Through the System
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────────────────────┐
|
||||
│ SENTIMENT ENGINE VOCABULARY FLOW │
|
||||
└─────────────────────────────────────────────────────────────────────────────────┘
|
||||
|
||||
RAW TEXT INPUT
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────────────┐
|
||||
│ ENTITY EXTRACTION (EntityExtractor) │
|
||||
│ • Ticker regex: \$?[A-Za-z]{2,10}\b │
|
||||
│ • Contract regex: 0x[a-fA-F0-9]{40} | base58 │
|
||||
│ • Alias lookup: Layer B (asset_aliases.yaml) + Layer C (known_entities) │
|
||||
│ • NER (spaCy): ORG, PRODUCT, GPE, PERSON → fuzzy map to tickers │
|
||||
│ Output: List[EntityExtraction{asset_id, mention_span, confidence, type}] │
|
||||
└─────────────────────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────────────┐
|
||||
│ SENTIMENT & EMOTION ANALYSIS (SentimentEmotionAnalyzer) │
|
||||
│ • FinBERT (ONNX/PyTorch/Mock) → [neg, neu, pos] probs │
|
||||
│ • CryptoSentimentCalibrator.calibrate(text, probs) ← LAYER A KEYWORDS │
|
||||
│ - _get_crypto_signal() uses CRYPTO_BULLISH/BEARISH_KEYWORDS │
|
||||
│ - WHALE_*_PHRASES weighted 5× │
|
||||
│ - Word-boundary regex for standard, substring for whale phrases │
|
||||
│ • Emotion model (DistilRoBERTa) → 6-class emotions │
|
||||
│ • Heuristic fallback if models unavailable │
|
||||
│ Output: SentimentScores(polarity, confidence, pos/neg/neu), EmotionScores │
|
||||
└─────────────────────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────────────┐
|
||||
│ EVENT CLASSIFICATION (EventClassifier) │
|
||||
│ • BERT classifier → 12-class event type │
|
||||
│ • Uses Layer F label schema (EVENT_LABELS) │
|
||||
└─────────────────────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────────────┐
|
||||
│ CREDIBILITY SCORING (CredibilityScorer) │
|
||||
│ • Source base_credibility from Layer D (source_credibility.yaml) │
|
||||
│ • Cross-source corroboration (in-memory cache) │
|
||||
│ • Temporal decay (half-life 30 days) │
|
||||
│ Output: CredibilityScore(composite, components) │
|
||||
└─────────────────────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────────────┐
|
||||
│ SIGNAL PROCESSING (SignalProcessor) │
|
||||
│ • fear_state = f(neg_sentiment, fear_emotion, event_fear) │
|
||||
│ • greed_state = f(pos_sentiment, greed_emotion, event_greed) │
|
||||
│ • pump_score = f(greed, joy, pos_events, intensity) │
|
||||
│ • dump_score = f(fear, anger, neg_events, intensity) │
|
||||
│ • VelocityComputer → hype_velocity, pub_velocity │
|
||||
│ • TemporalDecay (Layer E scoring.halflife_minutes) │
|
||||
│ Output: AssetSentiment per asset │
|
||||
└─────────────────────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────────────┐
|
||||
│ CENTROID REFINEMENT (ScoringEngine._refine_with_centroids) ← LAYER E │
|
||||
│ • Embed combined entity+event text via e5-large-v2 │
|
||||
│ • Cosine similarity to 6 parameter centroids (Layer E .npy files) │
|
||||
│ • Blend: 70% signal, 30% centroid │
|
||||
└─────────────────────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────────────┐
|
||||
│ AGGREGATION (Aggregator) │
|
||||
│ • Asset → Industry (Layer C asset_industry_map.yaml) │
|
||||
│ • Industry → Market │
|
||||
│ • Decay at each level (asset 30m, industry 60m, market 120m half-life) │
|
||||
│ Output: SentimentOutput(market, industries, assets) │
|
||||
└─────────────────────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────────────┐
|
||||
│ TRADING INTEGRATION │
|
||||
│ • ACB signals: market_sentiment_state, fear_state, greed_state, │
|
||||
│ hype_velocity, aggregate_pump_risk │
|
||||
│ • Book health veto: pump_score > 75 │
|
||||
│ • AlphaExitV7: dump_score > 70, fear_state > 80 │
|
||||
└─────────────────────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Consistency & Centralization Analysis
|
||||
|
||||
### 4.1 Current State: DISJOINT
|
||||
|
||||
| Aspect | Status | Detail |
|
||||
|--------|--------|--------|
|
||||
| **Single source of truth** | ❌ No | 6 independent stores with different schemas |
|
||||
| **Unified ID space** | ❌ No | Keywords (strings), aliases (ticker→ticker), entities (ticker→metadata), sources (source_id), centroids (param name), labels (enum values) |
|
||||
| **Versioning** | Partial | Git for code (Layer A, F), file mtime for YAML (B, C, D), file mtime for .npy (E) |
|
||||
| **Audit trail** | Partial | Git for code; DuckDB audit log for source credibility (D); none for centroids (E) |
|
||||
| **Hot-reload** | Mixed | YAML (B, C, D): yes; Python constants (A, F): no; .npy (E): no |
|
||||
| **Validation** | Minimal | `vocab_test_cases.json` tests Layer A only; no cross-layer validation |
|
||||
|
||||
### 4.2 Duplication & Drift Risks
|
||||
|
||||
| Risk | Location | Example |
|
||||
|------|----------|---------|
|
||||
| **Keyword ↔ Label drift** | Layer A vs Layer F | `CRYPTO_BULLISH_KEYWORDS` contains "moon" but `LABELING_GUIDELINES` lists "moon" under BULLISH emoji — consistent now, but no enforcement |
|
||||
| **Alias ↔ Entity drift** | Layer B vs Layer C | `asset_aliases.yaml` has "VITALIK" → "ETH"; `known_entities.yaml` has ETH entry — if one updated without other, resolution breaks |
|
||||
| **Centroid ↔ Keyword drift** | Layer E vs Layer A | Centroids built from keywords (TODO) but currently random; if keywords change, centroids stale |
|
||||
| **Source credibility ↔ Event outcome** | Layer D vs Labeling | `false_positive` event outcome adjusts credibility but event labels from Layer F — no automated loop |
|
||||
|
||||
---
|
||||
|
||||
## 5. Recommendations for Centralization
|
||||
|
||||
### 5.1 Immediate (Low Effort)
|
||||
|
||||
1. **Single Vocabulary Registry** — Create `config/vocabulary.yaml` with:
|
||||
```yaml
|
||||
sentiment_keywords:
|
||||
bullish: [...]
|
||||
bearish: [...]
|
||||
whale_bullish: [...]
|
||||
whale_bearish: [...]
|
||||
asset_aliases: {...} # merge Layer B
|
||||
known_entities: {...} # merge Layer C
|
||||
source_credibility: [...] # merge Layer D
|
||||
labeling_schema: # mirror Layer F
|
||||
sentiment: [BEARISH, BULLISH, NEUTRAL]
|
||||
events: [...]
|
||||
emotions: [...]
|
||||
```
|
||||
|
||||
2. **Runtime Loader** — `VocabularyRegistry` class loading YAML + `.npy` centroids, exposing typed accessors.
|
||||
|
||||
3. **Validation Tests** — Cross-layer consistency checks:
|
||||
- Every alias target exists in known_entities
|
||||
- Every whale phrase keyword appears in corresponding bullish/bearish list
|
||||
- Centroid rebuild script reads from `vocabulary.yaml` keyword lists
|
||||
|
||||
### 5.2 Medium Term
|
||||
|
||||
4. **Centroid Auto-Rebuild** — On vocabulary change, trigger centroid recomputation via encoder.
|
||||
|
||||
5. **Provenance Tracking** — Add `source: "keyword_list" | "centroid" | "heuristic"` to every score component.
|
||||
|
||||
6. **A/B Testing Framework** — Compare keyword-only vs. centroid-only vs. blended scoring.
|
||||
|
||||
### 5.3 Long Term
|
||||
|
||||
7. **Learned Vocabulary** — Replace hard-coded lists with learned token importance (attention weights, SHAP values) from fine-tuned model.
|
||||
|
||||
8. **Semantic Versioning** — `vocabulary.yaml` with `version: "2.1.0"`, migration scripts for schema changes.
|
||||
|
||||
---
|
||||
|
||||
## 6. File Inventory (Absolute Paths)
|
||||
|
||||
| Layer | File | Lines | Size | Last Modified |
|
||||
|-------|------|-------|------|---------------|
|
||||
| A | `/mnt/dolphinng5_predict/sentiment_engine/src/sentiment_engine/nlp/sentiment_emotion.py` | ~1,776 | ~68 KB | 2026-07-xx |
|
||||
| B | `/mnt/dolphinng5_predict/sentiment_engine/config/asset_aliases.yaml` | ~60 | 1.1 KB | 2026-07-xx |
|
||||
| C | `/mnt/dolphinng5_predict/sentiment_engine/config/known_entities.yaml` | ~55 | 1.8 KB | 2026-07-xx |
|
||||
| D | `/mnt/dolphinng5_predict/sentiment_engine/config/source_credibility.yaml` | ~70 | 2.9 KB | 2026-07-xx |
|
||||
| E | `/mnt/dolphinng5_predict/sentiment_engine/config/centroids/*.npy` (6 files) | — | 3.1 KB each | 2026-07-xx |
|
||||
| F | `/mnt/dolphinng5_predict/sentiment_engine/labeling_pipeline.py` | ~1,000+ | ~48 KB | 2026-07-xx |
|
||||
| Config | `/mnt/dolphinng5_predict/sentiment_engine/config/settings.yaml` | ~180 | 7.8 KB | 2026-07-xx |
|
||||
| Test | `/mnt/dolphinng5_predict/sentiment_engine/vocab_test_cases.json` | ~2,000 | 47 KB | 2026-07-xx |
|
||||
|
||||
---
|
||||
|
||||
## 7. Keyword Counts (Layer A)
|
||||
|
||||
| List | Count (approx) | Unique Stems |
|
||||
|------|----------------|--------------|
|
||||
| `CRYPTO_BULLISH_KEYWORDS` | 1,200+ | ~400 |
|
||||
| `CRYPTO_BEARISH_KEYWORDS` | 1,200+ | ~400 |
|
||||
| `WHALE_BULLISH_PHRASES` | 80 | 80 |
|
||||
| `WHALE_BEARISH_PHRASES` | 120 | 120 |
|
||||
| **Total** | **~2,600** | **~1,000** |
|
||||
|
||||
*Note: High duplication in lists (many variants: "surge", "surges", "surged", "surgeing", "surgeing").*
|
||||
|
||||
---
|
||||
|
||||
## 8. Test Coverage (Layer A)
|
||||
|
||||
**File:** `vocab_test_cases.json` — 200+ test cases
|
||||
**Coverage:** Basic positive/negative, whale phrases, compound phrases, edge cases
|
||||
**Run:** `pytest tests/test_crypto_sentiment_calibrator.py` (if exists) or manual via `labeling_pipeline.py`
|
||||
|
||||
---
|
||||
|
||||
## 9. Open Questions / TODOs
|
||||
|
||||
1. **Centroid building** — `_build_centroids()` currently uses random vectors. Implement keyword-driven centroid construction per `SENTIMENT_SPEC_IMPLEMENT_GUIDE.md`.
|
||||
|
||||
2. **Whale phrase matching** — Currently uses simple substring (`phrase in text_lower`). Should use word-boundary regex for consistency with standard keywords.
|
||||
|
||||
3. **Compound phrase deduplication** — Lists contain both `"golden.cross"` and `"golden cross"`. Normalize to single representation.
|
||||
|
||||
4. **Multi-word n-gram storage** — No explicit n-gram store beyond compound phrases in keyword lists. Consider adding n-gram frequency tracking from corpus.
|
||||
|
||||
5. **Language support** — Only English (`supported_languages: ["en"]`). Keyword lists are English-only.
|
||||
|
||||
6. **Dynamic keyword weighting** — All keywords equal weight (1). Could learn weights from labeled data.
|
||||
|
||||
---
|
||||
|
||||
## 10. Appendices
|
||||
|
||||
### 10.1 Full Keyword List Excerpt (Layer A)
|
||||
|
||||
See `sentiment_emotion.py` lines 200–1400 for complete lists.
|
||||
|
||||
### 10.2 Centroid Rebuild Procedure (When Implemented)
|
||||
|
||||
```bash
|
||||
# 1. Update vocabulary.yaml with new keywords
|
||||
# 2. Run rebuild script
|
||||
python -m sentiment_engine.scripts.rebuild_centroids
|
||||
# 3. Verify .npy files updated
|
||||
# 4. Restart scoring engine
|
||||
```
|
||||
|
||||
### 10.3 Hot-Reload Procedures
|
||||
|
||||
| Layer | Command |
|
||||
|-------|---------|
|
||||
| B, C, D | `POST /admin/reload-catalogue` (if API exposed) or restart `CatalogueManager` |
|
||||
| E | Restart `ScoringEngine` (no hot-reload) |
|
||||
| A, F | Full container rebuild + deploy |
|
||||
|
||||
---
|
||||
|
||||
**End of Specification**
|
||||
Reference in New Issue
Block a user