749 lines
25 KiB
Markdown
749 lines
25 KiB
Markdown
|
|
# Complete Guide: Pretraining & Fine-Tuning for Crypto Sentiment Engine
|
|||
|
|
|
|||
|
|
> **Target**: Transform pre-trained models (FinBERT, DistilRoBERTa, BERT-base) into crypto-native models
|
|||
|
|
> **Scope**: Sentiment (3-class), Emotion (6-class), Event Classification (12-class), NER (crypto entities)
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 📚 Part 1: Pre-Existing Labeled Datasets (Ready to Use)
|
|||
|
|
|
|||
|
|
### 1.1 Sentiment (3-class: Bearish/Bullish/Neutral)
|
|||
|
|
|
|||
|
|
| Dataset | Size | Labels | Source | Access |
|
|||
|
|
|---------|------|--------|--------|--------|
|
|||
|
|
| **Twitter Financial News** | 11,932 | Bearish/Bullish/Neutral | Twitter API | `hf://zeroshot/twitter-financial-news-sentiment` |
|
|||
|
|
| **Financial PhraseBank** | 4,840 | Positive/Negative/Neutral | Financial reports | `hf://takala/financial_phrasebank` |
|
|||
|
|
| **FiQA Sentiment** | 1,000+ | Positive/Negative/Neutral | Financial QA | `hf://explodinggradients/fiqa` |
|
|||
|
|
| **Crypto Twitter Sentiment** | ~50K | Bullish/Bearish/Neutral | Crypto Twitter | `hf://crypto-sentiment/crypto-tweets` |
|
|||
|
|
| **CryptoSentiment (Kaggle)** | ~20K | Positive/Negative/Neutral | Reddit/Twitter | Manual download |
|
|||
|
|
|
|||
|
|
**Loading Code**:
|
|||
|
|
```python
|
|||
|
|
from datasets import load_dataset
|
|||
|
|
|
|||
|
|
# Twitter Financial News (11,932 samples, 3 classes)
|
|||
|
|
ds = load_dataset("zeroshot/twitter-financial-news-sentiment")
|
|||
|
|
# Labels: 0=Bearish, 1=Bullish, 2=Neutral
|
|||
|
|
|
|||
|
|
# Financial PhraseBank (4,840 samples, 3 classes)
|
|||
|
|
ds = load_dataset("financial_phrasebank", "sentences_allagree")
|
|||
|
|
# Labels: Positive, Negative, Neutral
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 1.2 Crypto-Specific Sentiment Datasets
|
|||
|
|
|
|||
|
|
| Dataset | Size | Platform | Labels | Source |
|
|||
|
|
|---------|------|----------|--------|--------|
|
|||
|
|
| **Crypto Twitter Sentiment** | ~50K tweets | Twitter | Bullish/Bearish/Neutral | `hf://sharifamit/crypto-sentiment` |
|
|||
|
|
| **Crypto Reddit Sentiment** | ~30K posts | Reddit | Positive/Negative/Neutral | `hf://cryptonlp/reddit-sentiment` |
|
|||
|
|
| **Crypto Fear & Greed Index** | Historical | Alternative.me | 0-100 scale | API / CSV |
|
|||
|
|
| **Bitcoin Tweets Sentiment** | ~200K | Twitter | Positive/Negative | `hf://bitcoin-tweets-sentiment` |
|
|||
|
|
|
|||
|
|
### 1.3 Event Classification (12-class)
|
|||
|
|
|
|||
|
|
**No large public dataset exists** — this is the main gap. Available resources:
|
|||
|
|
|
|||
|
|
| Resource | Type | Size | Notes |
|
|||
|
|
|----------|------|------|-------|
|
|||
|
|
| **FEDS (Financial Event Detection)** | ~5K | 8 event types | Academic |
|
|||
|
|
| **FinRED** | ~10K | Relation extraction | Some events |
|
|||
|
|
| **Fincausal** | ~5K | Causal events | Shared task |
|
|||
|
|
| **MLEC (Multi-Lingual Event)** | ~20K | 10+ languages | Some events |
|
|||
|
|
|
|||
|
|
**Action Required**: Build custom event dataset (see Section 3).
|
|||
|
|
|
|||
|
|
### 1.4 Emotion (6-class: joy/fear/anger/greed/sadness/neutral)
|
|||
|
|
|
|||
|
|
| Dataset | Size | Domain | Labels |
|
|||
|
|
|---------|------|--------|--------|
|
|||
|
|
| **GoEmotions** | 58K | Reddit | 27 emotions → map to 6 |
|
|||
|
|
| **SemEval 2018 Task 1** | 11K | Twitter | 11 emotions |
|
|||
|
|
| **Financial Emotion** | ~5K | Financial news | Custom |
|
|||
|
|
|
|||
|
|
**Mapping GoEmotions → 6-class**:
|
|||
|
|
```python
|
|||
|
|
EMOTION_MAP = {
|
|||
|
|
"joy": ["joy", "amusement", "excitement", "gratitude", "love", "optimism", "pride", "relief"],
|
|||
|
|
"fear": ["fear", "nervousness", "anxiety"],
|
|||
|
|
"anger": ["anger", "annoyance", "disapproval", "disgust"],
|
|||
|
|
"greed": ["desire", "greed", "optimism"], # map from desire/optimism
|
|||
|
|
"sadness": ["sadness", "disappointment", "grief", "remorse"],
|
|||
|
|
"neutral": ["neutral", "confusion", "curiosity", "realization", "surprise"]
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 1.5 NER - Crypto Entities
|
|||
|
|
|
|||
|
|
| Dataset | Size | Entity Types |
|
|||
|
|
|---------|------|--------------|
|
|||
|
|
| **CryptoNER** | ~5K | Ticker, Contract, Person, Protocol, Exchange |
|
|||
|
|
| **CoNLL-2003** | 20K | PER, ORG, LOC, MISC (general) |
|
|||
|
|
| **FinBERT-NER** | ~5K | Financial entities |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 🏗️ Part 2: Data Collection & Labeling Pipeline
|
|||
|
|
|
|||
|
|
### 2.1 Data Sources for Raw Text Collection
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# config/data_sources.yaml
|
|||
|
|
raw_sources:
|
|||
|
|
twitter:
|
|||
|
|
- query: "bitcoin OR btc OR ethereum OR eth OR solana OR sol OR defi OR nft"
|
|||
|
|
lang: "en"
|
|||
|
|
limit: 10000
|
|||
|
|
reddit:
|
|||
|
|
subreddits: ["bitcoin", "ethereum", "cryptocurrency", "defi", "ethtrader", "bitcoinmarkets"]
|
|||
|
|
limit: 5000
|
|||
|
|
news_rss:
|
|||
|
|
feeds: ["coindesk.com", "cointelegraph.com", "theblock.co", "decrypt.co"]
|
|||
|
|
telegram:
|
|||
|
|
channels: ["defi_alpha", "whale_alert", "defi_pulse"]
|
|||
|
|
github:
|
|||
|
|
repos: ["ethereum", "solana-labs", "bitcoin"]
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 2.2 Automated Labeling Pipeline (Weak Supervision)
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# labeling/weak_supervision.py
|
|||
|
|
from snorkel.labeling import labeling_function, PandasLFApplier, LFAnalysis
|
|||
|
|
from snorkel.labeling.model import LabelModel
|
|||
|
|
|
|||
|
|
# Define labeling functions (LFs) for sentiment
|
|||
|
|
@labeling_function()
|
|||
|
|
def lf_bullish_keywords(x):
|
|||
|
|
bullish = ["moon", "pump", "bullish", "surge", "rally", "breakout", "ath", "long"]
|
|||
|
|
return 1 if any(w in x.text.lower() for w in bullish) else -1
|
|||
|
|
|
|||
|
|
@labeling_function()
|
|||
|
|
def lf_bearish_keywords(x):
|
|||
|
|
bearish = ["crash", "dump", "bearish", "dump", "panic", "rekt", "short", "collapse"]
|
|||
|
|
return 0 if any(w in x.text.lower() for w in bearish) else -1
|
|||
|
|
|
|||
|
|
@labeling_function()
|
|||
|
|
def lf_technical_bullish(x):
|
|||
|
|
tech = ["golden cross", "bull flag", "breakout", "support hold", "higher high"]
|
|||
|
|
return 1 if any(w in x.text.lower() for w in tech) else -1
|
|||
|
|
|
|||
|
|
@labeling_function()
|
|||
|
|
def lf_technical_bearish(x):
|
|||
|
|
tech = ["death cross", "bear flag", "breakdown", "resistance", "lower high"]
|
|||
|
|
return 0 if any(w in x.text.lower() for w in tech) else -1
|
|||
|
|
|
|||
|
|
@labeling_function()
|
|||
|
|
def lf_fundamental_bullish(x):
|
|||
|
|
fund = ["institutional", "etf", "adoption", "treasury", "whale buying", "accumulation"]
|
|||
|
|
return 1 if any(w in x.text.lower() for w in fund) else -1
|
|||
|
|
|
|||
|
|
@labeling_function()
|
|||
|
|
def lf_fundamental_bearish(x):
|
|||
|
|
fund = ["regulation", "ban", "hack", "exploit", "rug pull", "sec lawsuit"]
|
|||
|
|
return 0 if any(w in x.text.lower() for w in fund) else -1
|
|||
|
|
|
|||
|
|
@labeling_function()
|
|||
|
|
def lf_emoji_bullish(x):
|
|||
|
|
return 1 if any(e in x.text for e in ["🚀", "📈", "💎", "🙌", "🌙"]) else -1
|
|||
|
|
|
|||
|
|
@labeling_function()
|
|||
|
|
def lf_emoji_bearish(x):
|
|||
|
|
return 0 if any(e in x.text for e in ["📉", "😭", "💀", "🩸", "🧻"]) else -1
|
|||
|
|
|
|||
|
|
# Event LFs
|
|||
|
|
@labeling_function()
|
|||
|
|
def lf_hack_event(x):
|
|||
|
|
hack = ["hack", "exploit", "drain", "stolen", "vulnerability", "compromised"]
|
|||
|
|
return 2 if any(w in x.text.lower() for w in hack) else -1 # HACK=2
|
|||
|
|
|
|||
|
|
@labeling_function()
|
|||
|
|
def lf_listing_event(x):
|
|||
|
|
listing = ["listing", "listed", "debut", "goes live", "trading starts"]
|
|||
|
|
return 3 if any(w in x.text.lower() for w in listing) else -1 # LISTING=3
|
|||
|
|
|
|||
|
|
@labeling_function()
|
|||
|
|
def lf_regulatory_event(x):
|
|||
|
|
reg = ["sec", "cftc", "regulation", "lawsuit", "regulation", "compliance"]
|
|||
|
|
return 4 if any(w in x.text.lower() for w in reg) else -1 # REGULATORY=4
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 2.3 Human Annotation Workflow
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# labeling/annotation_interface.py
|
|||
|
|
import streamlit as st
|
|||
|
|
from datasets import Dataset
|
|||
|
|
|
|||
|
|
ANNOTATION_GUIDELINES = """
|
|||
|
|
## Sentiment Labeling Guidelines
|
|||
|
|
|
|||
|
|
### Labels: Bearish (0) | Neutral (1) | Bullish (2)
|
|||
|
|
|
|||
|
|
**Bullish (2)**: Explicit positive price action expectation
|
|||
|
|
- "BTC to $100k", "bullish on ETH", "accumulating", "moon", "pump"
|
|||
|
|
- Technical: "golden cross", "breakout", "breakout confirmed"
|
|||
|
|
- Fundamental: "institutional adoption", "ETF approval", "whale accumulation"
|
|||
|
|
|
|||
|
|
**Bearish (0)**: Explicit negative price action expectation
|
|||
|
|
- "crash incoming", "dump it", "top is in", "shorting", "rekt"
|
|||
|
|
- Technical: "death cross", "breakdown", "lower high", "resistance rejected"
|
|||
|
|
- Fundamental: "SEC lawsuit", "exchange hack", "regulation ban"
|
|||
|
|
|
|||
|
|
**Neutral (1)**: No clear directional bias
|
|||
|
|
- "BTC at $50k", "market consolidating", "waiting for direction"
|
|||
|
|
- Factual reporting without opinion: "BTC at $50k, ETH at $3k"
|
|||
|
|
|
|||
|
|
## Event Labeling Guidelines
|
|||
|
|
|
|||
|
|
### 12 Event Types:
|
|||
|
|
1. LISTING - New exchange listing, token debut
|
|||
|
|
2. DELISTING - Removal from exchange
|
|||
|
|
3. HACK - Exploit, drain, theft, vulnerability
|
|||
|
|
4. REGULATORY - SEC, CFTC, lawsuits, regulation
|
|||
|
|
5. GOVERNANCE - DAO votes, proposals, treasury
|
|||
|
|
6. UPGRADE - Hard fork, mainnet launch, protocol upgrade
|
|||
|
|
7. PARTNERSHIP - Integration, collaboration, alliance
|
|||
|
|
8. EARNINGS - Revenue, profit, financial results
|
|||
|
|
9. MACRO - Fed, rates, CPI, GDP, employment
|
|||
|
|
10. LIQUIDATION - Margin calls, cascade, cascading liquidations
|
|||
|
|
11. WHALE - Large transfers, accumulation, distribution
|
|||
|
|
12. MANIPULATION - Wash trading, spoofing, pump & dump
|
|||
|
|
"""
|
|||
|
|
|
|||
|
|
def create_annotation_dataset(raw_texts, output_path):
|
|||
|
|
"""Create annotation-ready dataset"""
|
|||
|
|
data = []
|
|||
|
|
for i, text in enumerate(raw_texts):
|
|||
|
|
data.append({
|
|||
|
|
"id": f"sample_{i:06d}",
|
|||
|
|
"text": text,
|
|||
|
|
"sentiment": None, # To be filled by annotator
|
|||
|
|
"events": [], # List of event types
|
|||
|
|
"entities": [], # Asset mentions
|
|||
|
|
"notes": ""
|
|||
|
|
)
|
|||
|
|
Dataset.from_list(data).to_json(output_path)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 🏋️ Part 3: Model Fine-Tuning Procedures
|
|||
|
|
|
|||
|
|
### 3.1 FinBERT Fine-Tuning (Sentiment)
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# training/finetune_finbert_sentiment.py
|
|||
|
|
from transformers import (
|
|||
|
|
AutoTokenizer, AutoModelForSequenceClassification,
|
|||
|
|
TrainingArguments, Trainer, EarlyStoppingCallback
|
|||
|
|
)
|
|||
|
|
from datasets import load_dataset
|
|||
|
|
import torch
|
|||
|
|
import numpy as np
|
|||
|
|
from sklearn.metrics import accuracy_score, f1_score, classification_report
|
|||
|
|
|
|||
|
|
# 1. Load & prepare data
|
|||
|
|
dataset = load_dataset("zeroshot/twitter-financial-news-sentiment")
|
|||
|
|
|
|||
|
|
# Add crypto-specific data
|
|||
|
|
crypto_ds = load_dataset("sharifamit/crypto-sentiment")
|
|||
|
|
# Combine & balance
|
|||
|
|
combined = concatenate_datasets([dataset["train"], crypto_ds["train"]])
|
|||
|
|
|
|||
|
|
# 2. Tokenizer
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained("ProsusAI/finbert")
|
|||
|
|
|
|||
|
|
def tokenize(batch):
|
|||
|
|
return tokenizer(batch["text"], truncation=True, max_length=256, padding="max_length")
|
|||
|
|
|
|||
|
|
tokenized = combined.map(tokenize, batched=True)
|
|||
|
|
|
|||
|
|
# 3. Model
|
|||
|
|
model = AutoModelForSequenceClassification.from_pretrained(
|
|||
|
|
"ProsusAI/finbert",
|
|||
|
|
num_labels=3,
|
|||
|
|
id2label={0: "Bearish", 1: "Bullish", 2: "Neutral"},
|
|||
|
|
label2id={"Bearish": 0, "Bullish": 1, "Neutral": 2}
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# 4. Class weights for imbalance
|
|||
|
|
class_weights = compute_class_weight("balanced", classes=np.unique(train_labels), y=train_labels)
|
|||
|
|
class_weights = torch.tensor(class_weights, dtype=torch.float)
|
|||
|
|
|
|||
|
|
# 4. Training arguments
|
|||
|
|
training_args = TrainingArguments(
|
|||
|
|
output_dir="./models/finbert-crypto-sentiment",
|
|||
|
|
num_train_epochs=5,
|
|||
|
|
per_device_train_batch_size=32,
|
|||
|
|
per_device_eval_batch_size=64,
|
|||
|
|
warmup_steps=500,
|
|||
|
|
weight_decay=0.01,
|
|||
|
|
learning_rate=2e-5,
|
|||
|
|
lr_scheduler_type="cosine",
|
|||
|
|
evaluation_strategy="epoch",
|
|||
|
|
save_strategy="epoch",
|
|||
|
|
load_best_model_at_end=True,
|
|||
|
|
metric_for_best_model="f1_macro",
|
|||
|
|
greater_is_better=True,
|
|||
|
|
fp16=True,
|
|||
|
|
logging_steps=100,
|
|||
|
|
report_to="wandb",
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# 5. Custom trainer with weighted loss
|
|||
|
|
class WeightedTrainer(Trainer):
|
|||
|
|
def compute_loss(self, model, inputs, return_outputs=False):
|
|||
|
|
labels = inputs.pop("labels")
|
|||
|
|
outputs = model(**inputs)
|
|||
|
|
logits = outputs.logits
|
|||
|
|
loss_fct = torch.nn.CrossEntropyLoss(weight=class_weights.to(logits.device))
|
|||
|
|
loss = loss_fct(logits.view(-1, 3), labels.view(-1))
|
|||
|
|
return (loss, outputs) if return_outputs else loss
|
|||
|
|
|
|||
|
|
# 6. Metrics
|
|||
|
|
def compute_metrics(eval_pred):
|
|||
|
|
logits, labels = eval_pred
|
|||
|
|
preds = np.argmax(logits, axis=-1)
|
|||
|
|
return {
|
|||
|
|
"accuracy": accuracy_score(labels, preds),
|
|||
|
|
"f1_macro": f1_score(labels, preds, average="macro"),
|
|||
|
|
"f1_per_class": f1_score(labels, preds, average=None).tolist()
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
trainer = WeightedTrainer(
|
|||
|
|
model=model,
|
|||
|
|
args=training_args,
|
|||
|
|
train_dataset=tokenized["train"],
|
|||
|
|
eval_dataset=tokenized["validation"],
|
|||
|
|
tokenizer=tokenizer,
|
|||
|
|
compute_metrics=compute_metrics,
|
|||
|
|
callbacks=[EarlyStoppingCallback(early_stopping_patience=3)]
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
trainer.train()
|
|||
|
|
trainer.save_model("./models/finbert-crypto-sentiment-final")
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 3.2 DistilRoBERTa Fine-Tuning (Emotion)
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# training/finetune_distilroberta_emotion.py
|
|||
|
|
from transformers import AutoTokenizer, AutoModelForSequenceClassification
|
|||
|
|
from datasets import load_dataset
|
|||
|
|
import torch
|
|||
|
|
|
|||
|
|
# 1. Load GoEmotions + financial emotion mapping
|
|||
|
|
go_emotions = load_dataset("go_emotions", "raw")
|
|||
|
|
# Filter & map to 6 classes using EMOTION_MAP
|
|||
|
|
|
|||
|
|
# Add financial emotion data
|
|||
|
|
fin_emotion = load_dataset("financial_emotion") # if available
|
|||
|
|
|
|||
|
|
# 2. Model: DistilRoBERTa-base (82M params)
|
|||
|
|
model_name = "j-hartmann/emotion-english-distilroberta-base"
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained(model_name)
|
|||
|
|
|
|||
|
|
model = AutoModelForSequenceClassification.from_pretrained(
|
|||
|
|
model_name,
|
|||
|
|
num_labels=6,
|
|||
|
|
id2label={0: "joy", 1: "fear", 2: "anger", 3: "greed", 4: "sadness", 5: "neutral"},
|
|||
|
|
label2id={"joy": 0, "fear": 1, "anger": 2, "greed": 3, "sadness": 4, "neutral": 5}
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# Freeze first 4 layers, fine-tune last 2 + classifier
|
|||
|
|
for param in model.distilroberta.embeddings.parameters():
|
|||
|
|
param.requires_grad = False
|
|||
|
|
for layer in model.distilroberta.transformer.layer[:4]:
|
|||
|
|
for param in layer.parameters():
|
|||
|
|
param.requires_grad = False
|
|||
|
|
|
|||
|
|
# Training args - lower LR for fine-tuning
|
|||
|
|
training_args = TrainingArguments(
|
|||
|
|
output_dir="./models/distilroberta-crypto-emotion",
|
|||
|
|
num_train_epochs=3,
|
|||
|
|
per_device_train_batch_size=16,
|
|||
|
|
learning_rate=1e-5, # Lower for fine-tuning
|
|||
|
|
warmup_ratio=0.1,
|
|||
|
|
# ... same as sentiment
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# Use multi-label if emotions can co-occur
|
|||
|
|
def compute_metrics(eval_pred):
|
|||
|
|
logits, labels = eval_pred
|
|||
|
|
preds = (torch.sigmoid(torch.tensor(logits)) > 0.5).int()
|
|||
|
|
return {
|
|||
|
|
"f1_micro": f1_score(labels, preds, average="micro"),
|
|||
|
|
"f1_macro": f1_score(labels, preds, average="macro"),
|
|||
|
|
"roc_auc": roc_auc_score(labels, torch.sigmoid(torch.tensor(logits)), average="macro")
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 3.3 BERT-base Fine-Tuning (Event Classification - 12 classes)
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# training/finetune_bert_events.py
|
|||
|
|
from transformers import AutoTokenizer, AutoModelForSequenceClassification
|
|||
|
|
from datasets import Dataset
|
|||
|
|
import json
|
|||
|
|
|
|||
|
|
# 1. CREATE CUSTOM EVENT DATASET
|
|||
|
|
# Since no public dataset exists, build from:
|
|||
|
|
# - RSS feeds with manual annotation
|
|||
|
|
# - News APIs with event tags
|
|||
|
|
# - Manual annotation of 5,000+ samples
|
|||
|
|
|
|||
|
|
EVENT_LABELS = [
|
|||
|
|
"listing", "delisting", "hack", "regulatory", "governance",
|
|||
|
|
"upgrade", "partnership", "earnings", "macro",
|
|||
|
|
"liquidation", "whale", "manipulation"
|
|||
|
|
]
|
|||
|
|
|
|||
|
|
label2id = {label: i for i, label in enumerate(EVENT_LABELS)}
|
|||
|
|
id2label = {i: label for i, label in enumerate(EVENT_LABELS)}
|
|||
|
|
|
|||
|
|
# 3. Multi-label classification (events can co-occur)
|
|||
|
|
model = AutoModelForSequenceClassification.from_pretrained(
|
|||
|
|
"bert-base-uncased",
|
|||
|
|
num_labels=12,
|
|||
|
|
problem_type="multi_label_classification",
|
|||
|
|
id2label=id2label,
|
|||
|
|
label2id=label2id
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# Multi-label loss
|
|||
|
|
def compute_loss(model, inputs):
|
|||
|
|
labels = inputs.pop("labels").float() # [batch, 12] multi-hot
|
|||
|
|
outputs = model(**inputs)
|
|||
|
|
logits = outputs.logits
|
|||
|
|
loss_fct = torch.nn.BCEWithLogitsLoss()
|
|||
|
|
loss = loss_fct(logits, labels)
|
|||
|
|
return loss
|
|||
|
|
|
|||
|
|
# Training with class weights for rare events (hack, manipulation)
|
|||
|
|
pos_weight = compute_pos_weight(train_labels) # [12]
|
|||
|
|
loss_fct = torch.nn.BCEWithLogitsLoss(pos_weight=pos_weight.to(device))
|
|||
|
|
|
|||
|
|
training_args = TrainingArguments(
|
|||
|
|
output_dir="./models/bert-crypto-events",
|
|||
|
|
num_train_epochs=5,
|
|||
|
|
per_device_train_batch_size=16,
|
|||
|
|
learning_rate=2e-5,
|
|||
|
|
# ... same
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# Multi-label metrics
|
|||
|
|
def compute_metrics(eval_pred):
|
|||
|
|
logits, labels = eval_pred
|
|||
|
|
probs = torch.sigmoid(torch.tensor(logits))
|
|||
|
|
preds = (probs > 0.5).int()
|
|||
|
|
return {
|
|||
|
|
"f1_micro": f1_score(labels, preds, average="micro"),
|
|||
|
|
"f1_macro": f1_score(labels, preds, average="macro"),
|
|||
|
|
"f1_per_class": f1_score(labels, preds, average=None).tolist(),
|
|||
|
|
"roc_auc_macro": roc_auc_score(labels, probs, average="macro"),
|
|||
|
|
"precision_at_k": precision_at_k(preds, labels, k=3)
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 3.4 Crypto NER Fine-Tuning
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# training/finetune_crypto_ner.py
|
|||
|
|
from transformers import AutoTokenizer, AutoModelForTokenClassification
|
|||
|
|
from datasets import load_dataset
|
|||
|
|
|
|||
|
|
# 1. Use CryptoNER dataset or create from CoNLL + crypto entities
|
|||
|
|
# Format: tokens + NER tags (B-ORG, I-ORG, B-TICKER, I-TICKER, B-CONTRACT, etc.)
|
|||
|
|
|
|||
|
|
CRYPTO_ENTITIES = [
|
|||
|
|
"TICKER", # BTC, ETH, SOL
|
|||
|
|
"CONTRACT", # 0x..., Solana addresses
|
|||
|
|
"PROTOCOL", # Uniswap, Aave, Lido
|
|||
|
|
"EXCHANGE", # Binance, Coinbase, Coinbase
|
|||
|
|
"PERSON", # Vitalik, CZ, SBF
|
|||
|
|
"CHAIN", # Ethereum, Solana, Arbitrum
|
|||
|
|
"TOKEN_STD", # ERC-20, SPL, BEP-20
|
|||
|
|
]
|
|||
|
|
|
|||
|
|
tag2id = {"O": 0}
|
|||
|
|
for ent in CRYPTO_ENTITIES:
|
|||
|
|
tag2id[f"B-{ent}"] = len(tag2id)
|
|||
|
|
tag2id[f"I-{ent}"] = len(tag2id)
|
|||
|
|
id2tag = {v: k for k, v in tag2id.items()}
|
|||
|
|
|
|||
|
|
# 2. Model
|
|||
|
|
model = AutoModelForTokenClassification.from_pretrained(
|
|||
|
|
"bert-base-cased",
|
|||
|
|
num_labels=len(tag2id),
|
|||
|
|
id2label=id2tag,
|
|||
|
|
label2id=tag2id
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# 3. Token-level metrics
|
|||
|
|
def compute_metrics(eval_pred):
|
|||
|
|
logits, labels = eval_pred
|
|||
|
|
preds = np.argmax(logits, axis=-1)
|
|||
|
|
# Remove padding (-100)
|
|||
|
|
true_labels = [[id2tag[l] for l in label if l != -100] for label in labels]
|
|||
|
|
true_preds = [[id2tag[p] for p, l in zip(pred, label) if l != -100] for pred, label in zip(preds, labels)]
|
|||
|
|
|
|||
|
|
from seqeval.metrics import f1_score, precision_score, recall_score
|
|||
|
|
return {
|
|||
|
|
"f1": f1_score(true_labels, true_preds),
|
|||
|
|
"precision": precision_score(true_labels, true_preds),
|
|||
|
|
"recall": recall_score(true_labels, true_preds)
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 📊 Part 4: Export to ONNX (Production)
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# export/export_all.py
|
|||
|
|
from optimum.onnxruntime import ORTModelForSequenceClassification, ORTModelForTokenClassification
|
|||
|
|
from transformers import AutoTokenizer
|
|||
|
|
from pathlib import Path
|
|||
|
|
|
|||
|
|
MODELS = {
|
|||
|
|
"finbert-crypto-sentiment": {
|
|||
|
|
"task": "text-classification",
|
|||
|
|
"output": "models/onnx/finbert-crypto",
|
|||
|
|
},
|
|||
|
|
"distilroberta-crypto-emotion": {
|
|||
|
|
"task": "text-classification",
|
|||
|
|
"output": "models/onnx/distilroberta-crypto-emotion",
|
|||
|
|
},
|
|||
|
|
"bert-crypto-events": {
|
|||
|
|
"task": "text-classification",
|
|||
|
|
"output": "models/onnx/bert-crypto-events",
|
|||
|
|
},
|
|||
|
|
"bert-crypto-ner": {
|
|||
|
|
"task": "token-classification",
|
|||
|
|
"output": "models/onnx/bert-crypto-ner",
|
|||
|
|
},
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
for name, config in MODELS.items():
|
|||
|
|
print(f"Exporting {name}...")
|
|||
|
|
model = ORTModelForSequenceClassification.from_pretrained(
|
|||
|
|
f"./models/{name}",
|
|||
|
|
export=True,
|
|||
|
|
task=config["task"]
|
|||
|
|
)
|
|||
|
|
model.save_pretrained(config["output"])
|
|||
|
|
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained(f"./models/{name}")
|
|||
|
|
tokenizer.save_pretrained(config["output"])
|
|||
|
|
|
|||
|
|
# Quantize for production
|
|||
|
|
from optimum.onnxruntime import ORTOptimizer
|
|||
|
|
from optimum.onnxruntime.configuration import OptimizationConfig
|
|||
|
|
|
|||
|
|
optimizer = ORTOptimizer.from_pretrained(config["output"])
|
|||
|
|
opt_config = OptimizationConfig(optimization_level=99, optimize_for_gpu=False)
|
|||
|
|
optimizer.optimize(save_dir=Path(config["output"]) / "quantized", optimization_config=opt_config)
|
|||
|
|
print(f" ✅ {name} exported & quantized")
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 📋 Part 5: Labeling Project Management
|
|||
|
|
|
|||
|
|
### 5.1 Annotation Team Setup
|
|||
|
|
|
|||
|
|
```yaml
|
|||
|
|
# labeling/project_config.yaml
|
|||
|
|
project:
|
|||
|
|
name: "crypto-sentiment-labeling"
|
|||
|
|
tasks:
|
|||
|
|
- sentiment: {classes: 3, priority: "high", target: 20000}
|
|||
|
|
- events: {classes: 12, priority: "high", target: 10000}
|
|||
|
|
- emotion: {classes: 6, priority: "medium", target: 10000}
|
|||
|
|
- ner: {classes: 14, priority: "medium", target: 5000}
|
|||
|
|
|
|||
|
|
annotators:
|
|||
|
|
- {name: "annotator_1", expertise: "crypto-trading", tasks: ["sentiment", "events"]}
|
|||
|
|
- {name: "annotator_2", expertise: "defi", tasks: ["events", "ner"]}
|
|||
|
|
- {name: "annotator_3", expertise: "technical-analysis", tasks: ["sentiment", "emotion"]}
|
|||
|
|
|
|||
|
|
quality_control:
|
|||
|
|
gold_standard_ratio: 0.1
|
|||
|
|
agreement_threshold: 0.8
|
|||
|
|
adjudicator: "senior_analyst"
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 5.2 Inter-Annotator Agreement Targets
|
|||
|
|
|
|||
|
|
| Task | Krippendorff's α Target | Cohen's κ Target |
|
|||
|
|
|------|------------------------|------------------|
|
|||
|
|
| Sentiment (3-class) | ≥ 0.80 | ≥ 0.75 |
|
|||
|
|
| Events (12-class) | ≥ 0.70 | ≥ 0.65 |
|
|||
|
|
| Emotion (6-class) | ≥ 0.75 | ≥ 0.70 |
|
|||
|
|
| NER (14 tags) | ≥ 0.85 | ≥ 0.80 |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 📈 Part 6: Evaluation & Validation
|
|||
|
|
|
|||
|
|
### 6.1 Test Sets (Holdout)
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# evaluation/test_sets.py
|
|||
|
|
# Curated test sets - NEVER used in training
|
|||
|
|
|
|||
|
|
SENTIMENT_TEST = [
|
|||
|
|
# Clear bullish
|
|||
|
|
("BTC breaks $100k! New ATH!", "Bullish"),
|
|||
|
|
("ETH to $10k by EOY, accumulate now", "Bullish"),
|
|||
|
|
("Institutional inflows hit record high", "Bullish"),
|
|||
|
|
|
|||
|
|
# Clear bearish
|
|||
|
|
("BTC crashes 50% in hours", "Bearish"),
|
|||
|
|
("Exchange hacked, $100M stolen", "Bearish"),
|
|||
|
|
("SEC sues major exchange", "Bearish"),
|
|||
|
|
|
|||
|
|
# Neutral
|
|||
|
|
("BTC at $50k, ETH at $3k", "Neutral"),
|
|||
|
|
("Market consolidating in range", "Neutral"),
|
|||
|
|
]
|
|||
|
|
|
|||
|
|
EVENT_TEST = [
|
|||
|
|
("Binance lists new token XYZ", ["listing"]),
|
|||
|
|
("Coinbase delists XRP", ["delisting"]),
|
|||
|
|
("DeFi protocol hacked, $50M drained", ["hack"]),
|
|||
|
|
("SEC sues Coinbase", ["regulatory"]),
|
|||
|
|
("Ethereum Cancun upgrade live", ["upgrade"]),
|
|||
|
|
("Whale moves 50k BTC to Binance", ["whale"]),
|
|||
|
|
]
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 6.2 Continuous Evaluation Pipeline
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
# evaluation/continuous_eval.py
|
|||
|
|
import schedule
|
|||
|
|
import time
|
|||
|
|
from datetime import datetime
|
|||
|
|
|
|||
|
|
def run_evaluation_cycle():
|
|||
|
|
"""Run nightly evaluation on fresh data"""
|
|||
|
|
# 1. Fetch last 24h predictions
|
|||
|
|
# 2. Compare with market outcome (price change)
|
|||
|
|
# 3. Log metrics to wandb/MLflow
|
|||
|
|
# 4. Alert if metrics degrade
|
|||
|
|
|
|||
|
|
metrics = evaluate_recent_predictions()
|
|||
|
|
log_to_monitoring(metrics)
|
|||
|
|
|
|||
|
|
if metrics["f1_macro"] < 0.6:
|
|||
|
|
alert_team("Model performance degraded!")
|
|||
|
|
|
|||
|
|
# Schedule daily
|
|||
|
|
schedule.every().day.at("02:00").do(run_evaluation_cycle)
|
|||
|
|
|
|||
|
|
while True:
|
|||
|
|
schedule.run_pending()
|
|||
|
|
time.sleep(60)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 💰 Part 7: Cost & Timeline Estimates
|
|||
|
|
|
|||
|
|
### 7.1 Compute Requirements
|
|||
|
|
|
|||
|
|
| Model | Parameters | GPU (Fine-tune) | Time (A100) | Cost @ $2/hr |
|
|||
|
|
|-------|------------|-----------------|-------------|--------------|
|
|||
|
|
| FinBERT (110M) | 110M | 1x A100 40GB | ~2 hrs | ~$4 |
|
|||
|
|
| DistilRoBERTa (82M) | 82M | 1x A100 40GB | ~1.5 hrs | ~$3 |
|
|||
|
|
| BERT-base (110M) | 110M | 1x A100 40GB | ~3 hrs | ~$6 |
|
|||
|
|
| BERT-base NER | 110M | 1x A100 40GB | ~4 hrs | ~$8 |
|
|||
|
|
|
|||
|
|
**Total compute: ~$20-30** (single run)
|
|||
|
|
|
|||
|
|
### 7.2 Labeling Costs
|
|||
|
|
|
|||
|
|
| Task | Samples | Annotators | Time/annotator | Cost @ $25/hr |
|
|||
|
|
|------|---------|------------|----------------|---------------|
|
|||
|
|
| Sentiment (3-class) | 20,000 | 3 | ~40 hrs | $3,000 |
|
|||
|
|
| Events (12-class) | 10,000 | 2 | ~60 hrs | $3,000 |
|
|||
|
|
| Emotion (6-class) | 10,000 | 2 | ~40 hrs | $2,000 |
|
|||
|
|
| NER (14 tags) | 5,000 | 2 | ~50 hrs | $2,500 |
|
|||
|
|
| **Total** | **45,000** | | | **~$10,500** |
|
|||
|
|
|
|||
|
|
**Alternative**: Use weak supervision (Snorkel) to reduce to ~$2,000
|
|||
|
|
|
|||
|
|
### 7.3 Timeline
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
Week 1-2: Data collection & weak supervision setup
|
|||
|
|
Week 3-4: Human annotation (parallel)
|
|||
|
|
Week 5: Data cleaning, train/val/test splits
|
|||
|
|
Week 6: FinBERT sentiment fine-tuning
|
|||
|
|
Week 7: DistilRoBERTa emotion fine-tuning
|
|||
|
|
Week 8: BERT event classification fine-tuning
|
|||
|
|
Week 9: BERT NER fine-tuning
|
|||
|
|
Week 10: ONNX export, quantization, integration testing
|
|||
|
|
Week 11-12: Shadow deployment, A/B testing
|
|||
|
|
Week 12+: Full production deployment
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 🎯 Part 8: Quick Start (Minimum Viable)
|
|||
|
|
|
|||
|
|
If you need **working models THIS WEEK**:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
# 1. Use existing models with prompt engineering (no training)
|
|||
|
|
python -c "
|
|||
|
|
from tweetnlp import load_model
|
|||
|
|
sentiment = load_model('sentiment')
|
|||
|
|
emotion = load_model('emotion')
|
|||
|
|
# Already fine-tuned on Twitter, works OK for crypto
|
|||
|
|
"
|
|||
|
|
|
|||
|
|
# 2. Apply weak supervision (Snorkel) - 1 day
|
|||
|
|
pip install snorkel
|
|||
|
|
python labeling/weak_supervision.py
|
|||
|
|
|
|||
|
|
# 3. Fine-tune FinBERT only (highest impact) - 1 day
|
|||
|
|
python training/finetune_finbert_sentiment.py
|
|||
|
|
|
|||
|
|
# 4. Export to ONNX - 30 min
|
|||
|
|
python export/export_all.py
|
|||
|
|
|
|||
|
|
# Total: ~2.5 days to "good enough" models
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 🔗 Key Resources
|
|||
|
|
|
|||
|
|
| Resource | Link |
|
|||
|
|
|----------|------|
|
|||
|
|
| **Twitter Financial News** | https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment |
|
|||
|
|
| **Financial PhraseBank** | https://huggingface.co/datasets/financial_phrasebank |
|
|||
|
|
| **GoEmotions** | https://huggingface.co/datasets/go_emotions |
|
|||
|
|
| **TweetNLP** | https://github.com/cardiffnlp/tweetnlp |
|
|||
|
|
| **Snorkel Tutorial** | https://www.snorkel.org/use-cases/ |
|
|||
|
|
| **HuggingFace Fine-tuning** | https://huggingface.co/docs/transformers/training |
|
|||
|
|
| **ONNX Export** | https://huggingface.co/docs/optimum/exporters/onnxruntime |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 🎯 Summary: What You Need To Do
|
|||
|
|
|
|||
|
|
| Priority | Action | Effort | Impact |
|
|||
|
|
|----------|--------|--------|--------|
|
|||
|
|
| **P0** | Fine-tune FinBERT on crypto sentiment | 1 day | Fixes polarity inversion |
|
|||
|
|
| **P0** | Build event dataset + fine-tune BERT | 3 days | Enables real event signals |
|
|||
|
|
| **P1** | Add crypto aliases + spaCy patterns | 4 hrs | Fixes entity gaps |
|
|||
|
|
| **P1** | Fine-tune DistilRoBERTa emotion | 1 day | Better emotion signals |
|
|||
|
|
| **P2** | Fine-tune NER | 1 day | Better entity extraction |
|
|||
|
|
| **P2** | Continuous eval pipeline | 4 hrs | Production monitoring |
|
|||
|
|
|
|||
|
|
**Total for production-ready**: ~1 week of focused work
|
|||
|
|
**Total for "good enough"**: ~2 days (FinBERT only + weak supervision)
|