Files
sentiment-engine/sentiment_engine/PRETRAINING_GUIDE.md

749 lines
25 KiB
Markdown
Raw Permalink Normal View History

# Complete Guide: Pretraining & Fine-Tuning for Crypto Sentiment Engine
> **Target**: Transform pre-trained models (FinBERT, DistilRoBERTa, BERT-base) into crypto-native models
> **Scope**: Sentiment (3-class), Emotion (6-class), Event Classification (12-class), NER (crypto entities)
---
## 📚 Part 1: Pre-Existing Labeled Datasets (Ready to Use)
### 1.1 Sentiment (3-class: Bearish/Bullish/Neutral)
| Dataset | Size | Labels | Source | Access |
|---------|------|--------|--------|--------|
| **Twitter Financial News** | 11,932 | Bearish/Bullish/Neutral | Twitter API | `hf://zeroshot/twitter-financial-news-sentiment` |
| **Financial PhraseBank** | 4,840 | Positive/Negative/Neutral | Financial reports | `hf://takala/financial_phrasebank` |
| **FiQA Sentiment** | 1,000+ | Positive/Negative/Neutral | Financial QA | `hf://explodinggradients/fiqa` |
| **Crypto Twitter Sentiment** | ~50K | Bullish/Bearish/Neutral | Crypto Twitter | `hf://crypto-sentiment/crypto-tweets` |
| **CryptoSentiment (Kaggle)** | ~20K | Positive/Negative/Neutral | Reddit/Twitter | Manual download |
**Loading Code**:
```python
from datasets import load_dataset
# Twitter Financial News (11,932 samples, 3 classes)
ds = load_dataset("zeroshot/twitter-financial-news-sentiment")
# Labels: 0=Bearish, 1=Bullish, 2=Neutral
# Financial PhraseBank (4,840 samples, 3 classes)
ds = load_dataset("financial_phrasebank", "sentences_allagree")
# Labels: Positive, Negative, Neutral
```
### 1.2 Crypto-Specific Sentiment Datasets
| Dataset | Size | Platform | Labels | Source |
|---------|------|----------|--------|--------|
| **Crypto Twitter Sentiment** | ~50K tweets | Twitter | Bullish/Bearish/Neutral | `hf://sharifamit/crypto-sentiment` |
| **Crypto Reddit Sentiment** | ~30K posts | Reddit | Positive/Negative/Neutral | `hf://cryptonlp/reddit-sentiment` |
| **Crypto Fear & Greed Index** | Historical | Alternative.me | 0-100 scale | API / CSV |
| **Bitcoin Tweets Sentiment** | ~200K | Twitter | Positive/Negative | `hf://bitcoin-tweets-sentiment` |
### 1.3 Event Classification (12-class)
**No large public dataset exists** — this is the main gap. Available resources:
| Resource | Type | Size | Notes |
|----------|------|------|-------|
| **FEDS (Financial Event Detection)** | ~5K | 8 event types | Academic |
| **FinRED** | ~10K | Relation extraction | Some events |
| **Fincausal** | ~5K | Causal events | Shared task |
| **MLEC (Multi-Lingual Event)** | ~20K | 10+ languages | Some events |
**Action Required**: Build custom event dataset (see Section 3).
### 1.4 Emotion (6-class: joy/fear/anger/greed/sadness/neutral)
| Dataset | Size | Domain | Labels |
|---------|------|--------|--------|
| **GoEmotions** | 58K | Reddit | 27 emotions → map to 6 |
| **SemEval 2018 Task 1** | 11K | Twitter | 11 emotions |
| **Financial Emotion** | ~5K | Financial news | Custom |
**Mapping GoEmotions → 6-class**:
```python
EMOTION_MAP = {
"joy": ["joy", "amusement", "excitement", "gratitude", "love", "optimism", "pride", "relief"],
"fear": ["fear", "nervousness", "anxiety"],
"anger": ["anger", "annoyance", "disapproval", "disgust"],
"greed": ["desire", "greed", "optimism"], # map from desire/optimism
"sadness": ["sadness", "disappointment", "grief", "remorse"],
"neutral": ["neutral", "confusion", "curiosity", "realization", "surprise"]
}
```
### 1.5 NER - Crypto Entities
| Dataset | Size | Entity Types |
|---------|------|--------------|
| **CryptoNER** | ~5K | Ticker, Contract, Person, Protocol, Exchange |
| **CoNLL-2003** | 20K | PER, ORG, LOC, MISC (general) |
| **FinBERT-NER** | ~5K | Financial entities |
---
## 🏗️ Part 2: Data Collection & Labeling Pipeline
### 2.1 Data Sources for Raw Text Collection
```python
# config/data_sources.yaml
raw_sources:
twitter:
- query: "bitcoin OR btc OR ethereum OR eth OR solana OR sol OR defi OR nft"
lang: "en"
limit: 10000
reddit:
subreddits: ["bitcoin", "ethereum", "cryptocurrency", "defi", "ethtrader", "bitcoinmarkets"]
limit: 5000
news_rss:
feeds: ["coindesk.com", "cointelegraph.com", "theblock.co", "decrypt.co"]
telegram:
channels: ["defi_alpha", "whale_alert", "defi_pulse"]
github:
repos: ["ethereum", "solana-labs", "bitcoin"]
```
### 2.2 Automated Labeling Pipeline (Weak Supervision)
```python
# labeling/weak_supervision.py
from snorkel.labeling import labeling_function, PandasLFApplier, LFAnalysis
from snorkel.labeling.model import LabelModel
# Define labeling functions (LFs) for sentiment
@labeling_function()
def lf_bullish_keywords(x):
bullish = ["moon", "pump", "bullish", "surge", "rally", "breakout", "ath", "long"]
return 1 if any(w in x.text.lower() for w in bullish) else -1
@labeling_function()
def lf_bearish_keywords(x):
bearish = ["crash", "dump", "bearish", "dump", "panic", "rekt", "short", "collapse"]
return 0 if any(w in x.text.lower() for w in bearish) else -1
@labeling_function()
def lf_technical_bullish(x):
tech = ["golden cross", "bull flag", "breakout", "support hold", "higher high"]
return 1 if any(w in x.text.lower() for w in tech) else -1
@labeling_function()
def lf_technical_bearish(x):
tech = ["death cross", "bear flag", "breakdown", "resistance", "lower high"]
return 0 if any(w in x.text.lower() for w in tech) else -1
@labeling_function()
def lf_fundamental_bullish(x):
fund = ["institutional", "etf", "adoption", "treasury", "whale buying", "accumulation"]
return 1 if any(w in x.text.lower() for w in fund) else -1
@labeling_function()
def lf_fundamental_bearish(x):
fund = ["regulation", "ban", "hack", "exploit", "rug pull", "sec lawsuit"]
return 0 if any(w in x.text.lower() for w in fund) else -1
@labeling_function()
def lf_emoji_bullish(x):
return 1 if any(e in x.text for e in ["🚀", "📈", "💎", "🙌", "🌙"]) else -1
@labeling_function()
def lf_emoji_bearish(x):
return 0 if any(e in x.text for e in ["📉", "😭", "💀", "🩸", "🧻"]) else -1
# Event LFs
@labeling_function()
def lf_hack_event(x):
hack = ["hack", "exploit", "drain", "stolen", "vulnerability", "compromised"]
return 2 if any(w in x.text.lower() for w in hack) else -1 # HACK=2
@labeling_function()
def lf_listing_event(x):
listing = ["listing", "listed", "debut", "goes live", "trading starts"]
return 3 if any(w in x.text.lower() for w in listing) else -1 # LISTING=3
@labeling_function()
def lf_regulatory_event(x):
reg = ["sec", "cftc", "regulation", "lawsuit", "regulation", "compliance"]
return 4 if any(w in x.text.lower() for w in reg) else -1 # REGULATORY=4
```
### 2.3 Human Annotation Workflow
```python
# labeling/annotation_interface.py
import streamlit as st
from datasets import Dataset
ANNOTATION_GUIDELINES = """
## Sentiment Labeling Guidelines
### Labels: Bearish (0) | Neutral (1) | Bullish (2)
**Bullish (2)**: Explicit positive price action expectation
- "BTC to $100k", "bullish on ETH", "accumulating", "moon", "pump"
- Technical: "golden cross", "breakout", "breakout confirmed"
- Fundamental: "institutional adoption", "ETF approval", "whale accumulation"
**Bearish (0)**: Explicit negative price action expectation
- "crash incoming", "dump it", "top is in", "shorting", "rekt"
- Technical: "death cross", "breakdown", "lower high", "resistance rejected"
- Fundamental: "SEC lawsuit", "exchange hack", "regulation ban"
**Neutral (1)**: No clear directional bias
- "BTC at $50k", "market consolidating", "waiting for direction"
- Factual reporting without opinion: "BTC at $50k, ETH at $3k"
## Event Labeling Guidelines
### 12 Event Types:
1. LISTING - New exchange listing, token debut
2. DELISTING - Removal from exchange
3. HACK - Exploit, drain, theft, vulnerability
4. REGULATORY - SEC, CFTC, lawsuits, regulation
5. GOVERNANCE - DAO votes, proposals, treasury
6. UPGRADE - Hard fork, mainnet launch, protocol upgrade
7. PARTNERSHIP - Integration, collaboration, alliance
8. EARNINGS - Revenue, profit, financial results
9. MACRO - Fed, rates, CPI, GDP, employment
10. LIQUIDATION - Margin calls, cascade, cascading liquidations
11. WHALE - Large transfers, accumulation, distribution
12. MANIPULATION - Wash trading, spoofing, pump & dump
"""
def create_annotation_dataset(raw_texts, output_path):
"""Create annotation-ready dataset"""
data = []
for i, text in enumerate(raw_texts):
data.append({
"id": f"sample_{i:06d}",
"text": text,
"sentiment": None, # To be filled by annotator
"events": [], # List of event types
"entities": [], # Asset mentions
"notes": ""
)
Dataset.from_list(data).to_json(output_path)
```
---
## 🏋️ Part 3: Model Fine-Tuning Procedures
### 3.1 FinBERT Fine-Tuning (Sentiment)
```python
# training/finetune_finbert_sentiment.py
from transformers import (
AutoTokenizer, AutoModelForSequenceClassification,
TrainingArguments, Trainer, EarlyStoppingCallback
)
from datasets import load_dataset
import torch
import numpy as np
from sklearn.metrics import accuracy_score, f1_score, classification_report
# 1. Load & prepare data
dataset = load_dataset("zeroshot/twitter-financial-news-sentiment")
# Add crypto-specific data
crypto_ds = load_dataset("sharifamit/crypto-sentiment")
# Combine & balance
combined = concatenate_datasets([dataset["train"], crypto_ds["train"]])
# 2. Tokenizer
tokenizer = AutoTokenizer.from_pretrained("ProsusAI/finbert")
def tokenize(batch):
return tokenizer(batch["text"], truncation=True, max_length=256, padding="max_length")
tokenized = combined.map(tokenize, batched=True)
# 3. Model
model = AutoModelForSequenceClassification.from_pretrained(
"ProsusAI/finbert",
num_labels=3,
id2label={0: "Bearish", 1: "Bullish", 2: "Neutral"},
label2id={"Bearish": 0, "Bullish": 1, "Neutral": 2}
)
# 4. Class weights for imbalance
class_weights = compute_class_weight("balanced", classes=np.unique(train_labels), y=train_labels)
class_weights = torch.tensor(class_weights, dtype=torch.float)
# 4. Training arguments
training_args = TrainingArguments(
output_dir="./models/finbert-crypto-sentiment",
num_train_epochs=5,
per_device_train_batch_size=32,
per_device_eval_batch_size=64,
warmup_steps=500,
weight_decay=0.01,
learning_rate=2e-5,
lr_scheduler_type="cosine",
evaluation_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="f1_macro",
greater_is_better=True,
fp16=True,
logging_steps=100,
report_to="wandb",
)
# 5. Custom trainer with weighted loss
class WeightedTrainer(Trainer):
def compute_loss(self, model, inputs, return_outputs=False):
labels = inputs.pop("labels")
outputs = model(**inputs)
logits = outputs.logits
loss_fct = torch.nn.CrossEntropyLoss(weight=class_weights.to(logits.device))
loss = loss_fct(logits.view(-1, 3), labels.view(-1))
return (loss, outputs) if return_outputs else loss
# 6. Metrics
def compute_metrics(eval_pred):
logits, labels = eval_pred
preds = np.argmax(logits, axis=-1)
return {
"accuracy": accuracy_score(labels, preds),
"f1_macro": f1_score(labels, preds, average="macro"),
"f1_per_class": f1_score(labels, preds, average=None).tolist()
}
trainer = WeightedTrainer(
model=model,
args=training_args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["validation"],
tokenizer=tokenizer,
compute_metrics=compute_metrics,
callbacks=[EarlyStoppingCallback(early_stopping_patience=3)]
)
trainer.train()
trainer.save_model("./models/finbert-crypto-sentiment-final")
```
### 3.2 DistilRoBERTa Fine-Tuning (Emotion)
```python
# training/finetune_distilroberta_emotion.py
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from datasets import load_dataset
import torch
# 1. Load GoEmotions + financial emotion mapping
go_emotions = load_dataset("go_emotions", "raw")
# Filter & map to 6 classes using EMOTION_MAP
# Add financial emotion data
fin_emotion = load_dataset("financial_emotion") # if available
# 2. Model: DistilRoBERTa-base (82M params)
model_name = "j-hartmann/emotion-english-distilroberta-base"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(
model_name,
num_labels=6,
id2label={0: "joy", 1: "fear", 2: "anger", 3: "greed", 4: "sadness", 5: "neutral"},
label2id={"joy": 0, "fear": 1, "anger": 2, "greed": 3, "sadness": 4, "neutral": 5}
)
# Freeze first 4 layers, fine-tune last 2 + classifier
for param in model.distilroberta.embeddings.parameters():
param.requires_grad = False
for layer in model.distilroberta.transformer.layer[:4]:
for param in layer.parameters():
param.requires_grad = False
# Training args - lower LR for fine-tuning
training_args = TrainingArguments(
output_dir="./models/distilroberta-crypto-emotion",
num_train_epochs=3,
per_device_train_batch_size=16,
learning_rate=1e-5, # Lower for fine-tuning
warmup_ratio=0.1,
# ... same as sentiment
)
# Use multi-label if emotions can co-occur
def compute_metrics(eval_pred):
logits, labels = eval_pred
preds = (torch.sigmoid(torch.tensor(logits)) > 0.5).int()
return {
"f1_micro": f1_score(labels, preds, average="micro"),
"f1_macro": f1_score(labels, preds, average="macro"),
"roc_auc": roc_auc_score(labels, torch.sigmoid(torch.tensor(logits)), average="macro")
}
```
### 3.3 BERT-base Fine-Tuning (Event Classification - 12 classes)
```python
# training/finetune_bert_events.py
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from datasets import Dataset
import json
# 1. CREATE CUSTOM EVENT DATASET
# Since no public dataset exists, build from:
# - RSS feeds with manual annotation
# - News APIs with event tags
# - Manual annotation of 5,000+ samples
EVENT_LABELS = [
"listing", "delisting", "hack", "regulatory", "governance",
"upgrade", "partnership", "earnings", "macro",
"liquidation", "whale", "manipulation"
]
label2id = {label: i for i, label in enumerate(EVENT_LABELS)}
id2label = {i: label for i, label in enumerate(EVENT_LABELS)}
# 3. Multi-label classification (events can co-occur)
model = AutoModelForSequenceClassification.from_pretrained(
"bert-base-uncased",
num_labels=12,
problem_type="multi_label_classification",
id2label=id2label,
label2id=label2id
)
# Multi-label loss
def compute_loss(model, inputs):
labels = inputs.pop("labels").float() # [batch, 12] multi-hot
outputs = model(**inputs)
logits = outputs.logits
loss_fct = torch.nn.BCEWithLogitsLoss()
loss = loss_fct(logits, labels)
return loss
# Training with class weights for rare events (hack, manipulation)
pos_weight = compute_pos_weight(train_labels) # [12]
loss_fct = torch.nn.BCEWithLogitsLoss(pos_weight=pos_weight.to(device))
training_args = TrainingArguments(
output_dir="./models/bert-crypto-events",
num_train_epochs=5,
per_device_train_batch_size=16,
learning_rate=2e-5,
# ... same
)
# Multi-label metrics
def compute_metrics(eval_pred):
logits, labels = eval_pred
probs = torch.sigmoid(torch.tensor(logits))
preds = (probs > 0.5).int()
return {
"f1_micro": f1_score(labels, preds, average="micro"),
"f1_macro": f1_score(labels, preds, average="macro"),
"f1_per_class": f1_score(labels, preds, average=None).tolist(),
"roc_auc_macro": roc_auc_score(labels, probs, average="macro"),
"precision_at_k": precision_at_k(preds, labels, k=3)
}
```
### 3.4 Crypto NER Fine-Tuning
```python
# training/finetune_crypto_ner.py
from transformers import AutoTokenizer, AutoModelForTokenClassification
from datasets import load_dataset
# 1. Use CryptoNER dataset or create from CoNLL + crypto entities
# Format: tokens + NER tags (B-ORG, I-ORG, B-TICKER, I-TICKER, B-CONTRACT, etc.)
CRYPTO_ENTITIES = [
"TICKER", # BTC, ETH, SOL
"CONTRACT", # 0x..., Solana addresses
"PROTOCOL", # Uniswap, Aave, Lido
"EXCHANGE", # Binance, Coinbase, Coinbase
"PERSON", # Vitalik, CZ, SBF
"CHAIN", # Ethereum, Solana, Arbitrum
"TOKEN_STD", # ERC-20, SPL, BEP-20
]
tag2id = {"O": 0}
for ent in CRYPTO_ENTITIES:
tag2id[f"B-{ent}"] = len(tag2id)
tag2id[f"I-{ent}"] = len(tag2id)
id2tag = {v: k for k, v in tag2id.items()}
# 2. Model
model = AutoModelForTokenClassification.from_pretrained(
"bert-base-cased",
num_labels=len(tag2id),
id2label=id2tag,
label2id=tag2id
)
# 3. Token-level metrics
def compute_metrics(eval_pred):
logits, labels = eval_pred
preds = np.argmax(logits, axis=-1)
# Remove padding (-100)
true_labels = [[id2tag[l] for l in label if l != -100] for label in labels]
true_preds = [[id2tag[p] for p, l in zip(pred, label) if l != -100] for pred, label in zip(preds, labels)]
from seqeval.metrics import f1_score, precision_score, recall_score
return {
"f1": f1_score(true_labels, true_preds),
"precision": precision_score(true_labels, true_preds),
"recall": recall_score(true_labels, true_preds)
}
```
---
## 📊 Part 4: Export to ONNX (Production)
```python
# export/export_all.py
from optimum.onnxruntime import ORTModelForSequenceClassification, ORTModelForTokenClassification
from transformers import AutoTokenizer
from pathlib import Path
MODELS = {
"finbert-crypto-sentiment": {
"task": "text-classification",
"output": "models/onnx/finbert-crypto",
},
"distilroberta-crypto-emotion": {
"task": "text-classification",
"output": "models/onnx/distilroberta-crypto-emotion",
},
"bert-crypto-events": {
"task": "text-classification",
"output": "models/onnx/bert-crypto-events",
},
"bert-crypto-ner": {
"task": "token-classification",
"output": "models/onnx/bert-crypto-ner",
},
}
for name, config in MODELS.items():
print(f"Exporting {name}...")
model = ORTModelForSequenceClassification.from_pretrained(
f"./models/{name}",
export=True,
task=config["task"]
)
model.save_pretrained(config["output"])
tokenizer = AutoTokenizer.from_pretrained(f"./models/{name}")
tokenizer.save_pretrained(config["output"])
# Quantize for production
from optimum.onnxruntime import ORTOptimizer
from optimum.onnxruntime.configuration import OptimizationConfig
optimizer = ORTOptimizer.from_pretrained(config["output"])
opt_config = OptimizationConfig(optimization_level=99, optimize_for_gpu=False)
optimizer.optimize(save_dir=Path(config["output"]) / "quantized", optimization_config=opt_config)
print(f" ✅ {name} exported & quantized")
```
---
## 📋 Part 5: Labeling Project Management
### 5.1 Annotation Team Setup
```yaml
# labeling/project_config.yaml
project:
name: "crypto-sentiment-labeling"
tasks:
- sentiment: {classes: 3, priority: "high", target: 20000}
- events: {classes: 12, priority: "high", target: 10000}
- emotion: {classes: 6, priority: "medium", target: 10000}
- ner: {classes: 14, priority: "medium", target: 5000}
annotators:
- {name: "annotator_1", expertise: "crypto-trading", tasks: ["sentiment", "events"]}
- {name: "annotator_2", expertise: "defi", tasks: ["events", "ner"]}
- {name: "annotator_3", expertise: "technical-analysis", tasks: ["sentiment", "emotion"]}
quality_control:
gold_standard_ratio: 0.1
agreement_threshold: 0.8
adjudicator: "senior_analyst"
```
### 5.2 Inter-Annotator Agreement Targets
| Task | Krippendorff's α Target | Cohen's κ Target |
|------|------------------------|------------------|
| Sentiment (3-class) | ≥ 0.80 | ≥ 0.75 |
| Events (12-class) | ≥ 0.70 | ≥ 0.65 |
| Emotion (6-class) | ≥ 0.75 | ≥ 0.70 |
| NER (14 tags) | ≥ 0.85 | ≥ 0.80 |
---
## 📈 Part 6: Evaluation & Validation
### 6.1 Test Sets (Holdout)
```python
# evaluation/test_sets.py
# Curated test sets - NEVER used in training
SENTIMENT_TEST = [
# Clear bullish
("BTC breaks $100k! New ATH!", "Bullish"),
("ETH to $10k by EOY, accumulate now", "Bullish"),
("Institutional inflows hit record high", "Bullish"),
# Clear bearish
("BTC crashes 50% in hours", "Bearish"),
("Exchange hacked, $100M stolen", "Bearish"),
("SEC sues major exchange", "Bearish"),
# Neutral
("BTC at $50k, ETH at $3k", "Neutral"),
("Market consolidating in range", "Neutral"),
]
EVENT_TEST = [
("Binance lists new token XYZ", ["listing"]),
("Coinbase delists XRP", ["delisting"]),
("DeFi protocol hacked, $50M drained", ["hack"]),
("SEC sues Coinbase", ["regulatory"]),
("Ethereum Cancun upgrade live", ["upgrade"]),
("Whale moves 50k BTC to Binance", ["whale"]),
]
```
### 6.2 Continuous Evaluation Pipeline
```python
# evaluation/continuous_eval.py
import schedule
import time
from datetime import datetime
def run_evaluation_cycle():
"""Run nightly evaluation on fresh data"""
# 1. Fetch last 24h predictions
# 2. Compare with market outcome (price change)
# 3. Log metrics to wandb/MLflow
# 4. Alert if metrics degrade
metrics = evaluate_recent_predictions()
log_to_monitoring(metrics)
if metrics["f1_macro"] < 0.6:
alert_team("Model performance degraded!")
# Schedule daily
schedule.every().day.at("02:00").do(run_evaluation_cycle)
while True:
schedule.run_pending()
time.sleep(60)
```
---
## 💰 Part 7: Cost & Timeline Estimates
### 7.1 Compute Requirements
| Model | Parameters | GPU (Fine-tune) | Time (A100) | Cost @ $2/hr |
|-------|------------|-----------------|-------------|--------------|
| FinBERT (110M) | 110M | 1x A100 40GB | ~2 hrs | ~$4 |
| DistilRoBERTa (82M) | 82M | 1x A100 40GB | ~1.5 hrs | ~$3 |
| BERT-base (110M) | 110M | 1x A100 40GB | ~3 hrs | ~$6 |
| BERT-base NER | 110M | 1x A100 40GB | ~4 hrs | ~$8 |
**Total compute: ~$20-30** (single run)
### 7.2 Labeling Costs
| Task | Samples | Annotators | Time/annotator | Cost @ $25/hr |
|------|---------|------------|----------------|---------------|
| Sentiment (3-class) | 20,000 | 3 | ~40 hrs | $3,000 |
| Events (12-class) | 10,000 | 2 | ~60 hrs | $3,000 |
| Emotion (6-class) | 10,000 | 2 | ~40 hrs | $2,000 |
| NER (14 tags) | 5,000 | 2 | ~50 hrs | $2,500 |
| **Total** | **45,000** | | | **~$10,500** |
**Alternative**: Use weak supervision (Snorkel) to reduce to ~$2,000
### 7.3 Timeline
```
Week 1-2: Data collection & weak supervision setup
Week 3-4: Human annotation (parallel)
Week 5: Data cleaning, train/val/test splits
Week 6: FinBERT sentiment fine-tuning
Week 7: DistilRoBERTa emotion fine-tuning
Week 8: BERT event classification fine-tuning
Week 9: BERT NER fine-tuning
Week 10: ONNX export, quantization, integration testing
Week 11-12: Shadow deployment, A/B testing
Week 12+: Full production deployment
```
---
## 🎯 Part 8: Quick Start (Minimum Viable)
If you need **working models THIS WEEK**:
```bash
# 1. Use existing models with prompt engineering (no training)
python -c "
from tweetnlp import load_model
sentiment = load_model('sentiment')
emotion = load_model('emotion')
# Already fine-tuned on Twitter, works OK for crypto
"
# 2. Apply weak supervision (Snorkel) - 1 day
pip install snorkel
python labeling/weak_supervision.py
# 3. Fine-tune FinBERT only (highest impact) - 1 day
python training/finetune_finbert_sentiment.py
# 4. Export to ONNX - 30 min
python export/export_all.py
# Total: ~2.5 days to "good enough" models
```
---
## 🔗 Key Resources
| Resource | Link |
|----------|------|
| **Twitter Financial News** | https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment |
| **Financial PhraseBank** | https://huggingface.co/datasets/financial_phrasebank |
| **GoEmotions** | https://huggingface.co/datasets/go_emotions |
| **TweetNLP** | https://github.com/cardiffnlp/tweetnlp |
| **Snorkel Tutorial** | https://www.snorkel.org/use-cases/ |
| **HuggingFace Fine-tuning** | https://huggingface.co/docs/transformers/training |
| **ONNX Export** | https://huggingface.co/docs/optimum/exporters/onnxruntime |
---
## 🎯 Summary: What You Need To Do
| Priority | Action | Effort | Impact |
|----------|--------|--------|--------|
| **P0** | Fine-tune FinBERT on crypto sentiment | 1 day | Fixes polarity inversion |
| **P0** | Build event dataset + fine-tune BERT | 3 days | Enables real event signals |
| **P1** | Add crypto aliases + spaCy patterns | 4 hrs | Fixes entity gaps |
| **P1** | Fine-tune DistilRoBERTa emotion | 1 day | Better emotion signals |
| **P2** | Fine-tune NER | 1 day | Better entity extraction |
| **P2** | Continuous eval pipeline | 4 hrs | Production monitoring |
**Total for production-ready**: ~1 week of focused work
**Total for "good enough"**: ~2 days (FinBERT only + weak supervision)