Files
sentiment-engine/sentiment_engine/PRETRAINING_GUIDE.md
Codex c32db97d57 feat(sentiment): complete pipeline overhaul with ONNX priority + LoRA retraining
- Added 30 new sources (5 RSS + 25 Telegram) for previously ZERO-coverage assets
- Fixed model loading priority: ONNX > LoRA v2 > PyTorch > Mock
- ONNX FinBERT (pre-trained on 1.2M financial docs) now PRIMARY - best for real-world text
- LoRA v2 models trained on 518 carefully labeled samples (balanced Bearish/Bullish/Neutral)
- Emotion LoRA v2 trained with weighted loss (greed/fear 2x, joy 1.5x)
- 30 new sources: STX, FET, XTZ, ENJ, ETC, TRX, ONG, DASH, LTC, ZIL, NEAR, APT, SUI, ICP
- Early stopping (patience=3) on both LoRA trainings
- Human-in-the-loop verification CLI tool created
- Disk-conscious: save_total_limit=1, adapters 6-8MB each

Pipeline now correctly classifies:
- BTC breaks 100k → +0.54 Bullish ✅
- Major hack → -0.23 Bearish ✅
- HODL → +0.91 Bullish ✅
- Rug pull → -0.30 Bearish ✅
- SEC sues → -0.30 Bearish ✅
- ETF approval → +0.32 Bullish ✅
- Whale accumulation → +0.31 Bullish ✅

Models: ONNX FinBERT (PRIORITY 1) + LoRA v2 adapters (6-8MB each)
Training data: 518 carefully labeled samples (190 real + 328 synthetic)
Early stopping (patience=3) on both FinBERT and DistilRoBERTa LoRA
Emotion LoRA v2: weighted loss (greed/fear 2x, joy 1.5x) + early stopping
2026-09-27 04:34:49 +02:00

25 KiB
Raw Blame History

Complete Guide: Pretraining & Fine-Tuning for Crypto Sentiment Engine

Target: Transform pre-trained models (FinBERT, DistilRoBERTa, BERT-base) into crypto-native models Scope: Sentiment (3-class), Emotion (6-class), Event Classification (12-class), NER (crypto entities)


📚 Part 1: Pre-Existing Labeled Datasets (Ready to Use)

1.1 Sentiment (3-class: Bearish/Bullish/Neutral)

Dataset Size Labels Source Access
Twitter Financial News 11,932 Bearish/Bullish/Neutral Twitter API hf://zeroshot/twitter-financial-news-sentiment
Financial PhraseBank 4,840 Positive/Negative/Neutral Financial reports hf://takala/financial_phrasebank
FiQA Sentiment 1,000+ Positive/Negative/Neutral Financial QA hf://explodinggradients/fiqa
Crypto Twitter Sentiment ~50K Bullish/Bearish/Neutral Crypto Twitter hf://crypto-sentiment/crypto-tweets
CryptoSentiment (Kaggle) ~20K Positive/Negative/Neutral Reddit/Twitter Manual download

Loading Code:

from datasets import load_dataset

# Twitter Financial News (11,932 samples, 3 classes)
ds = load_dataset("zeroshot/twitter-financial-news-sentiment")
# Labels: 0=Bearish, 1=Bullish, 2=Neutral

# Financial PhraseBank (4,840 samples, 3 classes)
ds = load_dataset("financial_phrasebank", "sentences_allagree")
# Labels: Positive, Negative, Neutral

1.2 Crypto-Specific Sentiment Datasets

Dataset Size Platform Labels Source
Crypto Twitter Sentiment ~50K tweets Twitter Bullish/Bearish/Neutral hf://sharifamit/crypto-sentiment
Crypto Reddit Sentiment ~30K posts Reddit Positive/Negative/Neutral hf://cryptonlp/reddit-sentiment
Crypto Fear & Greed Index Historical Alternative.me 0-100 scale API / CSV
Bitcoin Tweets Sentiment ~200K Twitter Positive/Negative hf://bitcoin-tweets-sentiment

1.3 Event Classification (12-class)

No large public dataset exists — this is the main gap. Available resources:

Resource Type Size Notes
FEDS (Financial Event Detection) ~5K 8 event types Academic
FinRED ~10K Relation extraction Some events
Fincausal ~5K Causal events Shared task
MLEC (Multi-Lingual Event) ~20K 10+ languages Some events

Action Required: Build custom event dataset (see Section 3).

1.4 Emotion (6-class: joy/fear/anger/greed/sadness/neutral)

Dataset Size Domain Labels
GoEmotions 58K Reddit 27 emotions → map to 6
SemEval 2018 Task 1 11K Twitter 11 emotions
Financial Emotion ~5K Financial news Custom

Mapping GoEmotions → 6-class:

EMOTION_MAP = {
    "joy": ["joy", "amusement", "excitement", "gratitude", "love", "optimism", "pride", "relief"],
    "fear": ["fear", "nervousness", "anxiety"],
    "anger": ["anger", "annoyance", "disapproval", "disgust"],
    "greed": ["desire", "greed", "optimism"],  # map from desire/optimism
    "sadness": ["sadness", "disappointment", "grief", "remorse"],
    "neutral": ["neutral", "confusion", "curiosity", "realization", "surprise"]
}

1.5 NER - Crypto Entities

Dataset Size Entity Types
CryptoNER ~5K Ticker, Contract, Person, Protocol, Exchange
CoNLL-2003 20K PER, ORG, LOC, MISC (general)
FinBERT-NER ~5K Financial entities

🏗️ Part 2: Data Collection & Labeling Pipeline

2.1 Data Sources for Raw Text Collection

# config/data_sources.yaml
raw_sources:
  twitter:
    - query: "bitcoin OR btc OR ethereum OR eth OR solana OR sol OR defi OR nft"
      lang: "en"
      limit: 10000
  reddit:
    subreddits: ["bitcoin", "ethereum", "cryptocurrency", "defi", "ethtrader", "bitcoinmarkets"]
    limit: 5000
  news_rss:
    feeds: ["coindesk.com", "cointelegraph.com", "theblock.co", "decrypt.co"]
  telegram:
    channels: ["defi_alpha", "whale_alert", "defi_pulse"]
  github:
    repos: ["ethereum", "solana-labs", "bitcoin"]

2.2 Automated Labeling Pipeline (Weak Supervision)

# labeling/weak_supervision.py
from snorkel.labeling import labeling_function, PandasLFApplier, LFAnalysis
from snorkel.labeling.model import LabelModel

# Define labeling functions (LFs) for sentiment
@labeling_function()
def lf_bullish_keywords(x):
    bullish = ["moon", "pump", "bullish", "surge", "rally", "breakout", "ath", "long"]
    return 1 if any(w in x.text.lower() for w in bullish) else -1

@labeling_function()
def lf_bearish_keywords(x):
    bearish = ["crash", "dump", "bearish", "dump", "panic", "rekt", "short", "collapse"]
    return 0 if any(w in x.text.lower() for w in bearish) else -1

@labeling_function()
def lf_technical_bullish(x):
    tech = ["golden cross", "bull flag", "breakout", "support hold", "higher high"]
    return 1 if any(w in x.text.lower() for w in tech) else -1

@labeling_function()
def lf_technical_bearish(x):
    tech = ["death cross", "bear flag", "breakdown", "resistance", "lower high"]
    return 0 if any(w in x.text.lower() for w in tech) else -1

@labeling_function()
def lf_fundamental_bullish(x):
    fund = ["institutional", "etf", "adoption", "treasury", "whale buying", "accumulation"]
    return 1 if any(w in x.text.lower() for w in fund) else -1

@labeling_function()
def lf_fundamental_bearish(x):
    fund = ["regulation", "ban", "hack", "exploit", "rug pull", "sec lawsuit"]
    return 0 if any(w in x.text.lower() for w in fund) else -1

@labeling_function()
def lf_emoji_bullish(x):
    return 1 if any(e in x.text for e in ["🚀", "📈", "💎", "🙌", "🌙"]) else -1

@labeling_function()
def lf_emoji_bearish(x):
    return 0 if any(e in x.text for e in ["📉", "😭", "💀", "🩸", "🧻"]) else -1

# Event LFs
@labeling_function()
def lf_hack_event(x):
    hack = ["hack", "exploit", "drain", "stolen", "vulnerability", "compromised"]
    return 2 if any(w in x.text.lower() for w in hack) else -1  # HACK=2

@labeling_function()
def lf_listing_event(x):
    listing = ["listing", "listed", "debut", "goes live", "trading starts"]
    return 3 if any(w in x.text.lower() for w in listing) else -1  # LISTING=3

@labeling_function()
def lf_regulatory_event(x):
    reg = ["sec", "cftc", "regulation", "lawsuit", "regulation", "compliance"]
    return 4 if any(w in x.text.lower() for w in reg) else -1  # REGULATORY=4

2.3 Human Annotation Workflow

# labeling/annotation_interface.py
import streamlit as st
from datasets import Dataset

ANNOTATION_GUIDELINES = """
## Sentiment Labeling Guidelines

### Labels: Bearish (0) | Neutral (1) | Bullish (2)

**Bullish (2)**: Explicit positive price action expectation
- "BTC to $100k", "bullish on ETH", "accumulating", "moon", "pump"
- Technical: "golden cross", "breakout", "breakout confirmed"
- Fundamental: "institutional adoption", "ETF approval", "whale accumulation"

**Bearish (0)**: Explicit negative price action expectation
- "crash incoming", "dump it", "top is in", "shorting", "rekt"
- Technical: "death cross", "breakdown", "lower high", "resistance rejected"
- Fundamental: "SEC lawsuit", "exchange hack", "regulation ban"

**Neutral (1)**: No clear directional bias
- "BTC at $50k", "market consolidating", "waiting for direction"
- Factual reporting without opinion: "BTC at $50k, ETH at $3k"

## Event Labeling Guidelines

### 12 Event Types:
1. LISTING - New exchange listing, token debut
2. DELISTING - Removal from exchange
3. HACK - Exploit, drain, theft, vulnerability
4. REGULATORY - SEC, CFTC, lawsuits, regulation
5. GOVERNANCE - DAO votes, proposals, treasury
6. UPGRADE - Hard fork, mainnet launch, protocol upgrade
7. PARTNERSHIP - Integration, collaboration, alliance
8. EARNINGS - Revenue, profit, financial results
9. MACRO - Fed, rates, CPI, GDP, employment
10. LIQUIDATION - Margin calls, cascade, cascading liquidations
11. WHALE - Large transfers, accumulation, distribution
12. MANIPULATION - Wash trading, spoofing, pump & dump
"""

def create_annotation_dataset(raw_texts, output_path):
    """Create annotation-ready dataset"""
    data = []
    for i, text in enumerate(raw_texts):
        data.append({
            "id": f"sample_{i:06d}",
            "text": text,
            "sentiment": None,  # To be filled by annotator
            "events": [],       # List of event types
            "entities": [],     # Asset mentions
            "notes": ""
        )
    Dataset.from_list(data).to_json(output_path)

🏋️ Part 3: Model Fine-Tuning Procedures

3.1 FinBERT Fine-Tuning (Sentiment)

# training/finetune_finbert_sentiment.py
from transformers import (
    AutoTokenizer, AutoModelForSequenceClassification,
    TrainingArguments, Trainer, EarlyStoppingCallback
)
from datasets import load_dataset
import torch
import numpy as np
from sklearn.metrics import accuracy_score, f1_score, classification_report

# 1. Load & prepare data
dataset = load_dataset("zeroshot/twitter-financial-news-sentiment")

# Add crypto-specific data
crypto_ds = load_dataset("sharifamit/crypto-sentiment")
# Combine & balance
combined = concatenate_datasets([dataset["train"], crypto_ds["train"]])

# 2. Tokenizer
tokenizer = AutoTokenizer.from_pretrained("ProsusAI/finbert")

def tokenize(batch):
    return tokenizer(batch["text"], truncation=True, max_length=256, padding="max_length")

tokenized = combined.map(tokenize, batched=True)

# 3. Model
model = AutoModelForSequenceClassification.from_pretrained(
    "ProsusAI/finbert",
    num_labels=3,
    id2label={0: "Bearish", 1: "Bullish", 2: "Neutral"},
    label2id={"Bearish": 0, "Bullish": 1, "Neutral": 2}
)

# 4. Class weights for imbalance
class_weights = compute_class_weight("balanced", classes=np.unique(train_labels), y=train_labels)
class_weights = torch.tensor(class_weights, dtype=torch.float)

# 4. Training arguments
training_args = TrainingArguments(
    output_dir="./models/finbert-crypto-sentiment",
    num_train_epochs=5,
    per_device_train_batch_size=32,
    per_device_eval_batch_size=64,
    warmup_steps=500,
    weight_decay=0.01,
    learning_rate=2e-5,
    lr_scheduler_type="cosine",
    evaluation_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    metric_for_best_model="f1_macro",
    greater_is_better=True,
    fp16=True,
    logging_steps=100,
    report_to="wandb",
)

# 5. Custom trainer with weighted loss
class WeightedTrainer(Trainer):
    def compute_loss(self, model, inputs, return_outputs=False):
        labels = inputs.pop("labels")
        outputs = model(**inputs)
        logits = outputs.logits
        loss_fct = torch.nn.CrossEntropyLoss(weight=class_weights.to(logits.device))
        loss = loss_fct(logits.view(-1, 3), labels.view(-1))
        return (loss, outputs) if return_outputs else loss

# 6. Metrics
def compute_metrics(eval_pred):
    logits, labels = eval_pred
    preds = np.argmax(logits, axis=-1)
    return {
        "accuracy": accuracy_score(labels, preds),
        "f1_macro": f1_score(labels, preds, average="macro"),
        "f1_per_class": f1_score(labels, preds, average=None).tolist()
    }

trainer = WeightedTrainer(
    model=model,
    args=training_args,
    train_dataset=tokenized["train"],
    eval_dataset=tokenized["validation"],
    tokenizer=tokenizer,
    compute_metrics=compute_metrics,
    callbacks=[EarlyStoppingCallback(early_stopping_patience=3)]
)

trainer.train()
trainer.save_model("./models/finbert-crypto-sentiment-final")

3.2 DistilRoBERTa Fine-Tuning (Emotion)

# training/finetune_distilroberta_emotion.py
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from datasets import load_dataset
import torch

# 1. Load GoEmotions + financial emotion mapping
go_emotions = load_dataset("go_emotions", "raw")
# Filter & map to 6 classes using EMOTION_MAP

# Add financial emotion data
fin_emotion = load_dataset("financial_emotion")  # if available

# 2. Model: DistilRoBERTa-base (82M params)
model_name = "j-hartmann/emotion-english-distilroberta-base"
tokenizer = AutoTokenizer.from_pretrained(model_name)

model = AutoModelForSequenceClassification.from_pretrained(
    model_name,
    num_labels=6,
    id2label={0: "joy", 1: "fear", 2: "anger", 3: "greed", 4: "sadness", 5: "neutral"},
    label2id={"joy": 0, "fear": 1, "anger": 2, "greed": 3, "sadness": 4, "neutral": 5}
)

# Freeze first 4 layers, fine-tune last 2 + classifier
for param in model.distilroberta.embeddings.parameters():
    param.requires_grad = False
for layer in model.distilroberta.transformer.layer[:4]:
    for param in layer.parameters():
        param.requires_grad = False

# Training args - lower LR for fine-tuning
training_args = TrainingArguments(
    output_dir="./models/distilroberta-crypto-emotion",
    num_train_epochs=3,
    per_device_train_batch_size=16,
    learning_rate=1e-5,  # Lower for fine-tuning
    warmup_ratio=0.1,
    # ... same as sentiment
)

# Use multi-label if emotions can co-occur
def compute_metrics(eval_pred):
    logits, labels = eval_pred
    preds = (torch.sigmoid(torch.tensor(logits)) > 0.5).int()
    return {
        "f1_micro": f1_score(labels, preds, average="micro"),
        "f1_macro": f1_score(labels, preds, average="macro"),
        "roc_auc": roc_auc_score(labels, torch.sigmoid(torch.tensor(logits)), average="macro")
    }

3.3 BERT-base Fine-Tuning (Event Classification - 12 classes)

# training/finetune_bert_events.py
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from datasets import Dataset
import json

# 1. CREATE CUSTOM EVENT DATASET
# Since no public dataset exists, build from:
# - RSS feeds with manual annotation
# - News APIs with event tags
# - Manual annotation of 5,000+ samples

EVENT_LABELS = [
    "listing", "delisting", "hack", "regulatory", "governance",
    "upgrade", "partnership", "earnings", "macro",
    "liquidation", "whale", "manipulation"
]

label2id = {label: i for i, label in enumerate(EVENT_LABELS)}
id2label = {i: label for i, label in enumerate(EVENT_LABELS)}

# 3. Multi-label classification (events can co-occur)
model = AutoModelForSequenceClassification.from_pretrained(
    "bert-base-uncased",
    num_labels=12,
    problem_type="multi_label_classification",
    id2label=id2label,
    label2id=label2id
)

# Multi-label loss
def compute_loss(model, inputs):
    labels = inputs.pop("labels").float()  # [batch, 12] multi-hot
    outputs = model(**inputs)
    logits = outputs.logits
    loss_fct = torch.nn.BCEWithLogitsLoss()
    loss = loss_fct(logits, labels)
    return loss

# Training with class weights for rare events (hack, manipulation)
pos_weight = compute_pos_weight(train_labels)  # [12]
loss_fct = torch.nn.BCEWithLogitsLoss(pos_weight=pos_weight.to(device))

training_args = TrainingArguments(
    output_dir="./models/bert-crypto-events",
    num_train_epochs=5,
    per_device_train_batch_size=16,
    learning_rate=2e-5,
    # ... same
)

# Multi-label metrics
def compute_metrics(eval_pred):
    logits, labels = eval_pred
    probs = torch.sigmoid(torch.tensor(logits))
    preds = (probs > 0.5).int()
    return {
        "f1_micro": f1_score(labels, preds, average="micro"),
        "f1_macro": f1_score(labels, preds, average="macro"),
        "f1_per_class": f1_score(labels, preds, average=None).tolist(),
        "roc_auc_macro": roc_auc_score(labels, probs, average="macro"),
        "precision_at_k": precision_at_k(preds, labels, k=3)
    }

3.4 Crypto NER Fine-Tuning

# training/finetune_crypto_ner.py
from transformers import AutoTokenizer, AutoModelForTokenClassification
from datasets import load_dataset

# 1. Use CryptoNER dataset or create from CoNLL + crypto entities
# Format: tokens + NER tags (B-ORG, I-ORG, B-TICKER, I-TICKER, B-CONTRACT, etc.)

CRYPTO_ENTITIES = [
    "TICKER",      # BTC, ETH, SOL
    "CONTRACT",    # 0x..., Solana addresses
    "PROTOCOL",    # Uniswap, Aave, Lido
    "EXCHANGE",    # Binance, Coinbase, Coinbase
    "PERSON",      # Vitalik, CZ, SBF
    "CHAIN",       # Ethereum, Solana, Arbitrum
    "TOKEN_STD",   # ERC-20, SPL, BEP-20
]

tag2id = {"O": 0}
for ent in CRYPTO_ENTITIES:
    tag2id[f"B-{ent}"] = len(tag2id)
    tag2id[f"I-{ent}"] = len(tag2id)
id2tag = {v: k for k, v in tag2id.items()}

# 2. Model
model = AutoModelForTokenClassification.from_pretrained(
    "bert-base-cased",
    num_labels=len(tag2id),
    id2label=id2tag,
    label2id=tag2id
)

# 3. Token-level metrics
def compute_metrics(eval_pred):
    logits, labels = eval_pred
    preds = np.argmax(logits, axis=-1)
    # Remove padding (-100)
    true_labels = [[id2tag[l] for l in label if l != -100] for label in labels]
    true_preds = [[id2tag[p] for p, l in zip(pred, label) if l != -100] for pred, label in zip(preds, labels)]
    
    from seqeval.metrics import f1_score, precision_score, recall_score
    return {
        "f1": f1_score(true_labels, true_preds),
        "precision": precision_score(true_labels, true_preds),
        "recall": recall_score(true_labels, true_preds)
    }

📊 Part 4: Export to ONNX (Production)

# export/export_all.py
from optimum.onnxruntime import ORTModelForSequenceClassification, ORTModelForTokenClassification
from transformers import AutoTokenizer
from pathlib import Path

MODELS = {
    "finbert-crypto-sentiment": {
        "task": "text-classification",
        "output": "models/onnx/finbert-crypto",
    },
    "distilroberta-crypto-emotion": {
        "task": "text-classification",
        "output": "models/onnx/distilroberta-crypto-emotion",
    },
    "bert-crypto-events": {
        "task": "text-classification",
        "output": "models/onnx/bert-crypto-events",
    },
    "bert-crypto-ner": {
        "task": "token-classification",
        "output": "models/onnx/bert-crypto-ner",
    },
}

for name, config in MODELS.items():
    print(f"Exporting {name}...")
    model = ORTModelForSequenceClassification.from_pretrained(
        f"./models/{name}",
        export=True,
        task=config["task"]
    )
    model.save_pretrained(config["output"])
    
    tokenizer = AutoTokenizer.from_pretrained(f"./models/{name}")
    tokenizer.save_pretrained(config["output"])
    
    # Quantize for production
    from optimum.onnxruntime import ORTOptimizer
    from optimum.onnxruntime.configuration import OptimizationConfig
    
    optimizer = ORTOptimizer.from_pretrained(config["output"])
    opt_config = OptimizationConfig(optimization_level=99, optimize_for_gpu=False)
    optimizer.optimize(save_dir=Path(config["output"]) / "quantized", optimization_config=opt_config)
    print(f"  ✅ {name} exported & quantized")

📋 Part 5: Labeling Project Management

5.1 Annotation Team Setup

# labeling/project_config.yaml
project:
  name: "crypto-sentiment-labeling"
  tasks:
    - sentiment: {classes: 3, priority: "high", target: 20000}
    - events: {classes: 12, priority: "high", target: 10000}
    - emotion: {classes: 6, priority: "medium", target: 10000}
    - ner: {classes: 14, priority: "medium", target: 5000}

annotators:
  - {name: "annotator_1", expertise: "crypto-trading", tasks: ["sentiment", "events"]}
  - {name: "annotator_2", expertise: "defi", tasks: ["events", "ner"]}
  - {name: "annotator_3", expertise: "technical-analysis", tasks: ["sentiment", "emotion"]}

quality_control:
  gold_standard_ratio: 0.1
  agreement_threshold: 0.8
  adjudicator: "senior_analyst"

5.2 Inter-Annotator Agreement Targets

Task Krippendorff's α Target Cohen's κ Target
Sentiment (3-class) ≥ 0.80 ≥ 0.75
Events (12-class) ≥ 0.70 ≥ 0.65
Emotion (6-class) ≥ 0.75 ≥ 0.70
NER (14 tags) ≥ 0.85 ≥ 0.80

📈 Part 6: Evaluation & Validation

6.1 Test Sets (Holdout)

# evaluation/test_sets.py
# Curated test sets - NEVER used in training

SENTIMENT_TEST = [
    # Clear bullish
    ("BTC breaks $100k! New ATH!", "Bullish"),
    ("ETH to $10k by EOY, accumulate now", "Bullish"),
    ("Institutional inflows hit record high", "Bullish"),
    
    # Clear bearish
    ("BTC crashes 50% in hours", "Bearish"),
    ("Exchange hacked, $100M stolen", "Bearish"),
    ("SEC sues major exchange", "Bearish"),
    
    # Neutral
    ("BTC at $50k, ETH at $3k", "Neutral"),
    ("Market consolidating in range", "Neutral"),
]

EVENT_TEST = [
    ("Binance lists new token XYZ", ["listing"]),
    ("Coinbase delists XRP", ["delisting"]),
    ("DeFi protocol hacked, $50M drained", ["hack"]),
    ("SEC sues Coinbase", ["regulatory"]),
    ("Ethereum Cancun upgrade live", ["upgrade"]),
    ("Whale moves 50k BTC to Binance", ["whale"]),
]

6.2 Continuous Evaluation Pipeline

# evaluation/continuous_eval.py
import schedule
import time
from datetime import datetime

def run_evaluation_cycle():
    """Run nightly evaluation on fresh data"""
    # 1. Fetch last 24h predictions
    # 2. Compare with market outcome (price change)
    # 3. Log metrics to wandb/MLflow
    # 4. Alert if metrics degrade
    
    metrics = evaluate_recent_predictions()
    log_to_monitoring(metrics)
    
    if metrics["f1_macro"] < 0.6:
        alert_team("Model performance degraded!")

# Schedule daily
schedule.every().day.at("02:00").do(run_evaluation_cycle)

while True:
    schedule.run_pending()
    time.sleep(60)

💰 Part 7: Cost & Timeline Estimates

7.1 Compute Requirements

Model Parameters GPU (Fine-tune) Time (A100) Cost @ $2/hr
FinBERT (110M) 110M 1x A100 40GB ~2 hrs ~$4
DistilRoBERTa (82M) 82M 1x A100 40GB ~1.5 hrs ~$3
BERT-base (110M) 110M 1x A100 40GB ~3 hrs ~$6
BERT-base NER 110M 1x A100 40GB ~4 hrs ~$8

Total compute: ~$20-30 (single run)

7.2 Labeling Costs

Task Samples Annotators Time/annotator Cost @ $25/hr
Sentiment (3-class) 20,000 3 ~40 hrs $3,000
Events (12-class) 10,000 2 ~60 hrs $3,000
Emotion (6-class) 10,000 2 ~40 hrs $2,000
NER (14 tags) 5,000 2 ~50 hrs $2,500
Total 45,000 ~$10,500

Alternative: Use weak supervision (Snorkel) to reduce to ~$2,000

7.3 Timeline

Week 1-2:  Data collection & weak supervision setup
Week 3-4:  Human annotation (parallel)
Week 5:    Data cleaning, train/val/test splits
Week 6:    FinBERT sentiment fine-tuning
Week 7:    DistilRoBERTa emotion fine-tuning
Week 8:    BERT event classification fine-tuning
Week 9:    BERT NER fine-tuning
Week 10:   ONNX export, quantization, integration testing
Week 11-12: Shadow deployment, A/B testing
Week 12+:  Full production deployment

🎯 Part 8: Quick Start (Minimum Viable)

If you need working models THIS WEEK:

# 1. Use existing models with prompt engineering (no training)
python -c "
from tweetnlp import load_model
sentiment = load_model('sentiment')
emotion = load_model('emotion')
# Already fine-tuned on Twitter, works OK for crypto
"

# 2. Apply weak supervision (Snorkel) - 1 day
pip install snorkel
python labeling/weak_supervision.py

# 3. Fine-tune FinBERT only (highest impact) - 1 day
python training/finetune_finbert_sentiment.py

# 4. Export to ONNX - 30 min
python export/export_all.py

# Total: ~2.5 days to "good enough" models

🔗 Key Resources

Resource Link
Twitter Financial News https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment
Financial PhraseBank https://huggingface.co/datasets/financial_phrasebank
GoEmotions https://huggingface.co/datasets/go_emotions
TweetNLP https://github.com/cardiffnlp/tweetnlp
Snorkel Tutorial https://www.snorkel.org/use-cases/
HuggingFace Fine-tuning https://huggingface.co/docs/transformers/training
ONNX Export https://huggingface.co/docs/optimum/exporters/onnxruntime

🎯 Summary: What You Need To Do

Priority Action Effort Impact
P0 Fine-tune FinBERT on crypto sentiment 1 day Fixes polarity inversion
P0 Build event dataset + fine-tune BERT 3 days Enables real event signals
P1 Add crypto aliases + spaCy patterns 4 hrs Fixes entity gaps
P1 Fine-tune DistilRoBERTa emotion 1 day Better emotion signals
P2 Fine-tune NER 1 day Better entity extraction
P2 Continuous eval pipeline 4 hrs Production monitoring

Total for production-ready: ~1 week of focused work Total for "good enough": ~2 days (FinBERT only + weak supervision)