Large Language Models (LLMs)
232 tools
232 tools · page 10 of 10
Humanloop helps teams evaluate, monitor, and iterate on LLM-powered features using both automated metrics and human feedback, aimed at improving reliability of production AI features.
Langfuse is an open-source platform for tracing, evaluating, and analyzing LLM application behavior, aimed at teams who want self-hostable observability tooling for their AI stack.
LangSmith traces, evaluates, and monitors language model application runs, aimed at developers who need visibility into why an LLM app produced a particular output in production.
LiteLLM lets developers call over a hundred different LLM APIs using a single, OpenAI-compatible format, simplifying the process of switching or comparing model providers.
LMQL is a programming language that lets developers write structured, constrained queries against large language models, aimed at more reliable and controllable LLM outputs than plain prompting.
OpenPipe captures a team's existing LLM API calls and uses that data to fine-tune smaller, cheaper models that match the original model's performance for that specific task.
Portkey acts as a gateway between an application and many different LLM providers, adding caching, fallbacks, and observability without requiring separate integration code for each model.
AI orchestration platform combining 9 frontier LLMs for verifiable responses
TensorZero unifies an LLM gateway, observability, and optimization tooling into one open-source framework, aimed at teams building and iterating on production-grade LLM applications.
Traceloop provides tracing and quality monitoring for LLM applications built on OpenTelemetry standards, aimed at teams that want observability tooling that isn't locked to one vendor.
Unstructured extracts and cleans text from PDFs, slides, and other document formats so the content can be fed into LLM and RAG pipelines, handling a common early bottleneck in AI projects.
Vellum provides tools for prompt engineering, workflow orchestration, and evaluation of LLM-powered features, aimed at product teams shipping AI features into existing applications.
Weights & Biases tracks machine learning experiments, model training runs, and LLM evaluations, aimed at AI teams that need to compare results across many model versions.
WhyLabs monitors machine learning and LLM applications in production for data quality issues and performance drift, aimed at ML teams catching problems before they affect users.