## Time Series Evals with OpenAI o1-preview

We benchmarked o1-preview on our hardest eval task - **time series trend evaluations**. This post compares that performance against GPT-4o-mini, Claude 3.5 sonnet, and GPT-4o.

[**o1-preview Time Series Evaluations Arize AI**](/content/blog/o1-preview-time-series-evaluations/index.html)

## Prompt Caching Benchmarking

We compare the performance and cost savings of **prompt caching on Anthropic vs OpenAI**.

[**How to Make Your AI App Feel Magical: Prompt Caching Arize AI**](/content/blog/prompt-caching-analysis/index.html)

## Multi-Agent Systems: Swarm

We compare and contrast **OpenAI’s experimental Swarm repo** against other popular multi-agent frameworks: **Autogen and CrewAI**.

[**Comparing OpenAI Swarm with other Multi Agent Frameworks Arize AI**](/content/blog/comparing-openai-swarm/index.html)

## Instrumenting LLMs with OTel

Lessons learned from our journey to one million downloads of our OpenTelemetry wrapper, [OpenInference](https://github.com/Arize-ai/openinference).

[**Zero to a Million: Instrumenting LLMs with OTEL Arize AI**](/content/blog/zero-to-a-million-instrumenting-llms-with-otel/index.html)

## Comparing Agent Frameworks

We built the [same agent](https://github.com/Arize-ai/phoenix/tree/main/examples/agent_framework_comparison) in LangGraph, LlamaIndex Workflows, CrewAI, Autogen, and pure code. See how each framework compares.

[**Comparing Agent Frameworks Arize AI**](/content/blog-course/llm-agent-how-to-set-up/comparing-agent-frameworks/index.html)

## Testing Generation in RAG

Testing the generation stage of RAG across GPT-4 and Claude 2.1.

[**Evaluating the Generation Stage in RAG Arize AI**](/content/blog/evaluating-the-generation-stage-in-rag/index.html)
