You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
An interactive marimo notebook on ICICLE AI Tapis services, a hands-on RAG playground that shows every step, from embeddings and chunking to retrieval, grounded prompting and LLM-as-judge evals. Built for newcomers to RAG who want to see how each piece works.
Calibrated LLM-as-a-judge evaluation pipeline on AWS Bedrock Claude Sonnet 4.6 scores Claude Haiku 4.5 on AlpacaEval via DeepEval G-Eval with versioned rubrics, persisted chain-of-thought, and blind human calibration
This repository provides a solution for generating detailed and thoughtful questions based on workout plans and evaluating their quality using the G-Eval metric. It is designed to assist fitness enthusiasts, trainers, and developers working with structured workout plans in improving the clarity, relevance, and usability of their questions.
Ši repozitorija skirta bakalauro darbui, kurio tikslas - tirti ir įgyvendinti lietuviškų pasirenkamojo atsakymo klausimų (MCQ) generavimo bei vertinimo procesą naudojant LLM.
DeepEval — independent third-party profile of a public API surface, by API Evangelist. DeepEval is an open-source LLM evaluation framework — built and maintained by Confident AI — for testing and benchmarking large language model applications. It is structured like Pytest but specialized for LLM systems, providing 40+ research-backed metrics (G-Eva