CMAESTrainer.train() now accepts workers parameter and passes it to
evaluate_candidate(), enabling parallel episode evaluation during
actual training (not just in tests/benchmarks).
Benchmark result: ProcessPoolExecutor is optimal (4.76x speedup).
Ray is slower (0.36x) due to head init + plasma overhead for 90 scenarios.