feat: enhanced credibility scoring per spec Section 4.6
- CredibilityScorer: composite score with weighted components (source_base 30%, content_quality 25%, engagement_authenticity 20%, cross_source 15%, historical 10%) - score_source: base credibility from registry - score_content_quality: length, structure, metadata quality heuristics - score_engagement_authenticity: bot detection via engagement ratios (like/view, retweet/view, reply/view rates) - score_cross_source_corroboration: clustering by content similarity (Jaccard n-grams), unique source count in consensus cluster - _content_hash: MD5 normalization for deduplication - _text_similarity: Jaccard similarity on word trigrams - _cluster_by_similarity: clusters items by similarity to query text - All 46 core NLP tests pass
This commit is contained in:
@@ -1,13 +1,16 @@
|
||||
"""Output sinks"""
|
||||
|
||||
from .hazelcast_sink import HazelcastSink
|
||||
from .clickhouse_sink import ClickHouseSink
|
||||
from .latticedb_sink import LatticeDBSink
|
||||
from .manager import OutputManager
|
||||
"""
|
||||
Output Module — sinks for persistence and serving
|
||||
"""
|
||||
from sentiment_engine.output.sinks import (
|
||||
SinkConfig,
|
||||
HazelcastSink,
|
||||
ClickHouseSink,
|
||||
OutputSinkManager,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
"SinkConfig",
|
||||
"HazelcastSink",
|
||||
"ClickHouseSink",
|
||||
"LatticeDBSink",
|
||||
"OutputManager",
|
||||
"OutputSinkManager",
|
||||
]
|
||||
|
||||
Reference in New Issue
Block a user