InData Science CollectivebyAparna Dhinakaran·Oct 8, 2025Should I Use the Same Model for My Eval as My Agent?Examining Self-Evaluation Bias of Major LLMs
InData Science CollectivebyAparna Dhinakaran·Aug 20, 2025LLM-as-a-Judge: When to Use Reasoning, CoT, and ExplanationsA guide on what to use where — and why
InArize AIbyAparna Dhinakaran·Aug 1, 2025What Is Prompt Learning?Prompts, like models, should improve with feedback — not stay static.A response icon2A response icon2
InTDS ArchivebyAparna Dhinakaran·Sep 21, 2024Choosing Between LLM Agent FrameworksThe tradeoffs between building bespoke code-based agents and the major agent frameworks.A response icon25A response icon25
InTDS ArchivebyAparna Dhinakaran·Aug 30, 2024Navigating the New Types of LLM Agents and ArchitecturesThe failure of ReAct agents gives way to a new generation of agents — and possibilitiesA response icon9A response icon9
InTDS ArchivebyAparna Dhinakaran·Jul 31, 2024Evaluating SQL Generation with LLM as a JudgeResults point to a promising approach
InTDS ArchivebyAparna Dhinakaran·May 1, 2024Large Language Model Performance in Time Series AnalysisHow do major LLMs stack up at detecting anomalies or movements in the data when given a large set of time series data within the context…A response icon3A response icon3
InTDS ArchivebyAparna Dhinakaran·Apr 6, 2024Tips for Getting the Generation Part Right in Retrieval Augmented GenerationResults from experiments to evaluate and compare GPT-4, Claude 2.1, and Claude 3.0 Opus
InTDS ArchivebyAparna Dhinakaran·Mar 26, 2024Model Evaluations Versus Task EvaluationsUnderstanding the difference for LLM applications
InTDS ArchivebyAparna Dhinakaran·Mar 8, 2024Why You Should Not Use Numeric Evals For LLM As a JudgeTesting major LLMs on how well they conduct numeric evaluations