Quantitative performance benchmarking of machine learning classifiers for phishing detection, utilizing precision, recall, and F1-score optimization on high-dimensional labeled datasets.
-
Updated
May 28, 2026 - Python
Quantitative performance benchmarking of machine learning classifiers for phishing detection, utilizing precision, recall, and F1-score optimization on high-dimensional labeled datasets.
Statistical evaluation of Gemma refusal robustness using direction ablation and official SORRY-Bench scoring.
frontier-evals-harness is a lightweight framework for benchmarking frontier language models. It provides deterministic suite versioning, modular adapters, standardized scoring, and paired statistical comparisons with confidence intervals. Built for regression tracking and analysis, it enables reproducible evaluation without infrastructure.
DTU 02445 individual assignment: subject-level evaluation of regression models on heart-rate data
Provider-neutral AI evaluation toolkit for reusable test cases, regression datasets, provider adapters, deterministic and LLM judges, agent trajectory and security evaluation, statistical analysis, quality/cost/latency trade-offs, online experiments, evidence guardrails, and release gates.
Classification models for detecting fake reviews and predicting software bugs. Includes implementations of decision trees, bagging, random forests, logistic regression, and Naive Bayes, with statistical evaluation using McNemar's test.
To associate your repository with the statistical-evaluation topic, visit your repo's landing page and select "manage topics."