advanced-evaluation
Implement production-grade LLM-as-a-judge pipelines for model evaluation, including pairwise comparison, direct scoring, bias mitigation, and rubric generation.
Explore AI agent skills related to model evaluation. Browse installable Claude Code and automation skills on Mentalok Skills Hub.
Discover reusable agent skills, browse implementation details, and find the right skill for your workflow.
3 skills found
Implement production-grade LLM-as-a-judge pipelines for model evaluation, including pairwise comparison, direct scoring, bias mitigation, and rubric generation.
Connect your AI agent to the Hugging Face Hub via MCP. Search models, datasets, and papers, manage repos, run cloud compute jobs, and invoke Gradio Spaces as functional AI tools.
Classical machine learning with scikit-learn. Use for classification, regression, clustering, dimensionality reduction, preprocessing, model evaluation, and building robust ML pipelines in Python.