blog
- agentic-ai
- ir
- llm
- math
- nlg-basics
- software
- visualization
•
•
•
•
•
•
-
Evaluating an Agentic AI Application 2/3: What Research and Our Experience Taught Us About LLM-as-a-Judge
An LLM judge is a measurement instrument, not an oracle. Research shows that judges carry systematic bias, that the judge prompt is part of the instrument, and that confident verdicts can flip under pressure. Here is what we changed in our own evaluation approach as a result.
-
Evaluating an Agentic AI Application 1/3: Summarizing my Experience Integrating DeepEval with Arize Phoenix at Codify
How can you make sure that your AI agents' output is truthful and faithful to your query? This is what inspired our journey integrating the evaluation platform DeepEval with the LLM trace platform Arize Phoenix. In this blog post, I will present the challenges I faced in this journey and how I solved them.
-
A Brief Introduction to (some of) the Math Behind the Transformer
The Transformer architecture, introduced in 2017 by Google, is the basis for every state-of-the-art generative AI model today. Understanding (some of) the math behind it is essential for understanding the Transformer itself and today's generative AI models.
-
A Brief Introduction to Natural Language Generation (NLG)
Natural language generation (NLG), i.e., the generation of text using language models and their causal language modeling capabilities, is not as black box as people make it out to be. In fact, a user can control many parameters. (Note that it is still basically like magic and that me and others are just trying our best)