- Parallel model initialization (asyncio.gather) reduces startup from 114s+ to ~70-80s
- Progress logging with timestamps at each stage for visibility
- Error handling with fallback retry on individual model failures
- Mock tokenizer fallback for transformers/tokenizers version conflicts
- Timeout-aware ONNX loading with separate tokenizer/model timing
All 3 ONNX models (FinBERT 417.9MB, bert-base-event 417.9MB, entity extractor) now load reliably.
End-to-end pipeline: entities → sentiment per asset → events → credibility working.