Files
sentiment-engine/sentiment_engine/fix_false_positives.py
Codex c32db97d57 feat(sentiment): complete pipeline overhaul with ONNX priority + LoRA retraining
- Added 30 new sources (5 RSS + 25 Telegram) for previously ZERO-coverage assets
- Fixed model loading priority: ONNX > LoRA v2 > PyTorch > Mock
- ONNX FinBERT (pre-trained on 1.2M financial docs) now PRIMARY - best for real-world text
- LoRA v2 models trained on 518 carefully labeled samples (balanced Bearish/Bullish/Neutral)
- Emotion LoRA v2 trained with weighted loss (greed/fear 2x, joy 1.5x)
- 30 new sources: STX, FET, XTZ, ENJ, ETC, TRX, ONG, DASH, LTC, ZIL, NEAR, APT, SUI, ICP
- Early stopping (patience=3) on both LoRA trainings
- Human-in-the-loop verification CLI tool created
- Disk-conscious: save_total_limit=1, adapters 6-8MB each

Pipeline now correctly classifies:
- BTC breaks 100k → +0.54 Bullish ✅
- Major hack → -0.23 Bearish ✅
- HODL → +0.91 Bullish ✅
- Rug pull → -0.30 Bearish ✅
- SEC sues → -0.30 Bearish ✅
- ETF approval → +0.32 Bullish ✅
- Whale accumulation → +0.31 Bullish ✅

Models: ONNX FinBERT (PRIORITY 1) + LoRA v2 adapters (6-8MB each)
Training data: 518 carefully labeled samples (190 real + 328 synthetic)
Early stopping (patience=3) on both FinBERT and DistilRoBERTa LoRA
Emotion LoRA v2: weighted loss (greed/fear 2x, joy 1.5x) + early stopping
2026-09-27 04:34:49 +02:00

28 lines
1.3 KiB
Python

with open('src/sentiment_engine/nlp/entity_extraction.py', 'r') as f:
lines = f.readlines()
new_lines = []
for line in lines:
stripped = line.strip()
if stripped == '"MOVING", "HARD", "SOFT", "FAST", "SLOW", "BIG", "SMALL",':
new_lines.append(' "MOVING", "HARD", "SOFT", "FAST", "SLOW", "BIG", "SMALL",\n')
elif stripped == '"LONG", "SHORT", "HIGH", "LOW", "OPEN", "CLOSE",':
new_lines.append(' "LONG", "SHORT", "HIGH", "LOW", "OPEN", "CLOSE",\n')
elif stripped == '"BULL", "BEAR", "FLAT", "VOL", "VOLS",':
new_lines.append(' "BULL", "BEAR", "FLAT", "VOL", "VOLS",\n')
elif stripped == '"BID", "ASK", "MID", "VWAP", "TWAP",':
new_lines.append(' "BID", "ASK", "MID", "VWAP", "TWAP",\n')
elif stripped == '"RSI", "MACD", "BB", "EMA", "SMA", "WMA",':
new_lines.append(' "RSI", "MACD", "BB", "EMA", "SMA", "WMA",\n')
elif stripped == '"ATR", "ADX", "CCI", "STOCH", "RSI",':
new_lines.append(' "ATR", "ADX", "CCI", "STOCH", "RSI",\n')
elif stripped == '"K", "M", "B", "T", "MM", "BB", "TT",':
new_lines.append(' "K", "M", "B", "T", "MM", "BB", "TT",\n')
else:
new_lines.append(line)
with open('src/sentiment_engine/nlp/entity_extraction.py', 'w') as f:
f.writelines(new_lines)
print('Fixed indentation')