Files
sentiment-engine/sentiment_engine/create_final_set.py
Codex c32db97d57 feat(sentiment): complete pipeline overhaul with ONNX priority + LoRA retraining
- Added 30 new sources (5 RSS + 25 Telegram) for previously ZERO-coverage assets
- Fixed model loading priority: ONNX > LoRA v2 > PyTorch > Mock
- ONNX FinBERT (pre-trained on 1.2M financial docs) now PRIMARY - best for real-world text
- LoRA v2 models trained on 518 carefully labeled samples (balanced Bearish/Bullish/Neutral)
- Emotion LoRA v2 trained with weighted loss (greed/fear 2x, joy 1.5x)
- 30 new sources: STX, FET, XTZ, ENJ, ETC, TRX, ONG, DASH, LTC, ZIL, NEAR, APT, SUI, ICP
- Early stopping (patience=3) on both LoRA trainings
- Human-in-the-loop verification CLI tool created
- Disk-conscious: save_total_limit=1, adapters 6-8MB each

Pipeline now correctly classifies:
- BTC breaks 100k → +0.54 Bullish ✅
- Major hack → -0.23 Bearish ✅
- HODL → +0.91 Bullish ✅
- Rug pull → -0.30 Bearish ✅
- SEC sues → -0.30 Bearish ✅
- ETF approval → +0.32 Bullish ✅
- Whale accumulation → +0.31 Bullish ✅

Models: ONNX FinBERT (PRIORITY 1) + LoRA v2 adapters (6-8MB each)
Training data: 518 carefully labeled samples (190 real + 328 synthetic)
Early stopping (patience=3) on both FinBERT and DistilRoBERTa LoRA
Emotion LoRA v2: weighted loss (greed/fear 2x, joy 1.5x) + early stopping
2026-09-27 04:34:49 +02:00

58 lines
1.7 KiB
Python

#!/usr/bin/env python3
"""
Create final high-quality labeled dataset.
"""
import json
from pathlib import Path
from collections import Counter
all_labeled = []
for fname in [
"labeled_verified.jsonl",
"labeled_expanded.jsonl",
"labeled_large.jsonl",
"labeled_output.jsonl",
"labeled_real_world.jsonl",
]:
path = Path(f"/mnt/dolphinng5_predict/sentiment_engine/data/{fname}")
if path.exists():
with open(path) as f:
for line in f:
try:
item = json.loads(line.strip())
labels = item.get("labels", {})
text = item.get("text", item.get("raw_text", ""))
if text and labels.get("sentiment"):
all_labeled.append({
"text": text,
"sentiment": labels["sentiment"],
"event_type": labels.get("event_type", "unknown"),
"entities": labels.get("entities", []),
"source": fname,
})
except Exception as e:
pass
# Deduplicate
seen = set()
unique = []
for d in all_labeled:
h = hash(d["text"][:200])
if h not in seen:
seen.add(h)
unique.append(d)
print(f"Total unique labeled samples: {len(unique)}")
sent_dist = Counter(d["sentiment"] for d in unique)
print(f"Sentiment distribution: {dict(sent_dist)}")
# Save
with open("/mnt/dolphinng5_predict/sentiment_engine/data/final_labeled_set.jsonl", "w") as f:
for d in unique:
json.dump(d, f)
f.write("\n")
print(f"\nSaved {len(unique)} samples to final_labeled_set.jsonl")