Best Arize AI Podcasts (2025)

1
The Illusion of Thinking: What the Apple AI Paper Says About LLM Reasoning 30:35

Play Pause

9h ago30:35

30:35

This week we discuss The Illusion of Thinking, a new paper from researchers at Apple that challenges today’s evaluation methods and introduces a new benchmark: synthetic puzzles with controllable complexity and clean logic. Their findings? Large Reasoning Models (LRMs) show surprising failure modes, including a complete collapse on high-complexity …

1
‘Failures across the board’: Dara Tarkowski on Synapse case 18:59

4d ago18:59

18:59

Tarkowski, managing partner at Actuate Law, shares a legal perspective on the lawsuit that Yotta, a savings app provider, filed against Evolve Bank & Trust, in September and recently amended.

1
Accurate KV Cache Quantization with Outlier Tokens Tracing 25:11

17d ago25:11

25:11

We discuss Accurate KV Cache Quantization with Outlier Tokens Tracing, a deep dive into improving the efficiency of LLM inference. The authors enhance KV Cache quantization, a technique for reducing memory and compute costs during inference, by introducing a method to identify and exclude outlier tokens that hurt quantization accuracy, striking a b…

1
Can banks beat the 80/20 rule for generative AI? 21:54

18d ago21:54

21:54

Sid Khosla, EY Americas banking and capital markets Leader, predicts that over the next two years, 20% of generative AI cases will drive 80% of the value across financial institutions. In this podcast, he explains what those use cases are and how banks can make the most of them.

1
What will AI-driven banking look like in the future? 23:34

1M ago23:34

23:34

Theo Lau, co-founder of Unconventional Ventures and author of the new book Banking on Artificial Intelligence, shares a vision of how banks could help consumers navigate financial uncertainties.

1
Scalable Chain of Thoughts via Elastic Reasoning 28:54

1M ago28:54

28:54

In this week's episode, we talk about Elastic Reasoning, a novel framework designed to enhance the efficiency and scalability of large reasoning models by explicitly separating the reasoning process into two distinct phases: thinking and solution. This separation allows for independent allocation of computational budgets, addressing challenges rela…

1
‘You can’t make it a side gig’: Brett Pharr on banking as a service 17:28

2M ago17:28

17:28

It takes $50 million to $100 million dollars, three to five years of losses and a complete commitment from the board to make it as a banking-as-a-service bank, the CEO of Pathward bank says in the latest American Banker podcast.

1
Sleep-time Compute: Beyond Inference Scaling at Test-time 30:24

2M ago30:24

30:24

What if your LLM could think ahead—preparing answers before questions are even asked? In this week's paper read, we dive into a groundbreaking new paper from researchers at Letta, introducing sleep-time compute: a novel technique that lets models do their heavy lifting offline, well before the user query arrives. By predicting likely questions and …

1
How J.P. Morgan is helping to build connected cars 22:01

2M ago22:01

22:01

Rob Abrams, CEO of Mobility Payment Solutions at J.P. Morgan Payments, is overseeing the development of in-car wallet systems that turn cars into rolling credit cards. He explains his vision of what connected cars could look like and do in the future.

1
LibreEval: The Largest Open Source Benchmark for RAG Hallucination Detection 27:19

2M ago27:19

27:19

For this week's paper read, we dive into our own research. We wanted to create a replicable, evolving dataset that can keep pace with model training so that you always know you're testing with data your model has never seen before. We also saw the prohibitively high cost of running LLM evals at scale, and have used our data to fine-tune a series of…

1
Banks struggle to keep up with threat of AI deepfakes 15:22

2M ago15:22

15:22

Valerie Abend, Accenture’s financial services cybersecurity lead, explains what banks get wrong about fending off AI-based threats and what they should do instead.

1
AI Benchmark Deep Dive: Gemini 2.5 and Humanity's Last Exam 26:11

3M ago26:11

26:11

This week we talk about modern AI benchmarks, taking a close look at Google's recent Gemini 2.5 release and its performance on key evaluations, notably Humanity's Last Exam (HLE). In the session we covered Gemini 2.5's architecture, its advancements in reasoning and multimodality, and its impressive context window. We also talked about how benchmar…

1
'It can eliminate things we don't like doing': Vlad Lukic on AI 18:54

3M ago18:54

18:54

Generative AI will remove toil from our day to day jobs, argues Lukic, who is managing director and senior partner at Boston Consulting Group.

1
Model Context Protocol (MCP) 15:03

3M ago15:03

15:03

We cover Anthropic’s groundbreaking Model Context Protocol (MCP). Though it was released in November 2024, we've been seeing a lot of hype around it lately, and thought it was well worth digging into. Learn how this open standard is revolutionizing AI by enabling seamless integration between LLMs and external data sources, fundamentally transformin…

1
Could industry standards prevent the next Synapse-style mess? 17:09

3M ago17:09

17:09

“We want to put banks in the risk management driver’s seat,” says Sima Gandhi, co-founder of the Council for Fintech Ecosystem Standards, which has worked with a group of fintechs to create risk and compliance standards banks can use to evaluate their fintech partners.

1
‘It’s a tremendous priority’: Huntington CFO Wasserman on AI 22:48

4M ago22:48

22:48

Software code generation and knowledge management are two of the places the bank has begun using generative AI to improve efficiency.

1
AI Roundup: DeepSeek’s Big Moves, Claude 3.7, and the Latest Breakthroughs 30:23

4M ago30:23

30:23

This week, we're mixing things up a little bit. Instead of diving deep into a single research paper, we cover the biggest AI developments from the past few weeks. We break down key announcements, including: DeepSeek’s Big Launch Week: A look at FlashMLA (DeepSeek’s new approach to efficient inference) and DeepEP (their enhanced pretraining method).…

1
How DeepSeek is Pushing the Boundaries of AI Development 29:54

4M ago29:54

29:54

This week, we dive into DeepSeek. SallyAnn DeLucia, Product Manager at Arize, and Nick Luzio, a Solutions Engineer, break down key insights on a model that have dominating headlines for its significant breakthrough in inference speed over other models. What’s next for AI (and open source)? From training strategies to real-world performance, here’s …

1
How one New York bank is getting younger people to save 17:07

4M ago17:07

17:07

Thomas Rudzewick, CEO of Maspeth Savings Bank worries about the growing crisis of low savings among millennials and Gen Z. He believes banks like his can help reverse this trend with financial literacy and innovative savings tools.

1
Multiagent Finetuning: A Conversation with Researcher Yilun Du 30:03

5M ago30:03

30:03

We talk to Google DeepMind Senior Research Scientist (and incoming Assistant Professor at Harvard), Yilun Du, about his latest paper, "Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains." This paper introduces a multiagent finetuning framework that enhances the performance and diversity of language models by employing a society o…

1
East West Bank’s CEO on how the bank coped with LA wildfires 31:04

5M ago31:04

31:04

The bank, which is headquartered in Pasadena, had to quickly switch to remote work for many employees and come up with relief programs for customers whose homes and businesses were destroyed by fire.

1
Where Citi Ventures is placing its fintech bets in 2025 31:20

5M ago31:20

31:20

Arvind Purushotham, head of Citi Ventures, shares where he and his team see opportunities and how they vet tech startups.

1
Training Large Language Models to Reason in Continuous Latent Space 24:58

5M ago24:58

24:58

LLMs have typically been restricted to reason in the "language space," where chain-of-thought (CoT) is used to solve complex reasoning problems. But a new paper argues that language space may not always be the best for reasoning. In this paper read, we cover an exciting new technique from a team at Meta called Chain of Continuous Thought—also known…

1
Why banks keep failing at money laundering 29:12

5M ago29:12

29:12

A lack of resources is one common cause of AML penalties, says consultant Aaron Ansari.

1
LLMs as Judges: A Comprehensive Survey on LLM-Based Evaluation Methods 28:57

6M ago28:57

28:57

We discuss a major survey of work and research on LLM-as-Judge from the last few years. "LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods" systematically examines the LLMs-as-Judge framework across five dimensions: functionality, methodology, applications, meta-evaluation, and limitations. This survey gives us a birds eye view…

1
The banks that implement AI well, from titans to mavericks 16:11

6M ago16:11

16:11

Some banks are “punching above their weight,” according to Dan Latimore, chief research officer at The Financial Revolutionist. Here’s how they do it.

1
Merge, Ensemble, and Cooperate! A Survey on Collaborative LLM Strategies 28:47

6M ago28:47

28:47

LLMs have revolutionized natural language processing, showcasing remarkable versatility and capabilities. But individual LLMs often exhibit distinct strengths and weaknesses, influenced by differences in their training corpora. This diversity poses a challenge: how can we maximize the efficiency and utility of LLMs? A new paper, "Merge, Ensemble, a…

1
The case for a human crime officer in every bank 20:57

7M ago20:57

20:57

Ian Mitchell, founder of The Noble, an organization that works with law enforcement and with banks to fight human crime trafficking, explains some of his group’s recent work and why banks need someone dedicated to human crime.

1
Agent-as-a-Judge: Evaluate Agents with Agents 24:54

7M ago24:54

24:54

This week, we break down the “Agent-as-a-Judge” framework—a new agent evaluation paradigm that’s kind of like getting robots to grade each other’s homework. Where typical evaluation methods focus solely on outcomes or demand extensive manual work, this approach uses agent systems to evaluate agent systems, offering intermediate feedback throughout …

1
AARP’s Jilenne Gunther has advice for banks on elder fraud 20:38

7M ago20:38

20:38

Elder financial exploitation has been a problem for banks for years, and it’s getting worse. Gunther offers practical suggestions for what banks should do when they suspect an older customer is a victim.

1
Introduction to OpenAI's Realtime API 29:56

7M ago29:56

29:56

We break down OpenAI’s realtime API. Learn how to seamlessly integrate powerful language models into your applications for instant, context-aware responses that drive user engagement. Whether you’re building chatbots, dynamic content tools, or enhancing real-time collaboration, we walk through the API’s capabilities, potential use cases, and best p…

1
Southern Bancorp’s answer to the home affordability crisis 19:43

8M ago19:43

19:43

The Arkansas community development financial institution has an ambitious goal of making $500 million worth of mortgages in rural, minority and low-income neighborhoods.

1
Swarm: OpenAI's Experimental Approach to Multi-Agent Systems 46:46

8M ago46:46

46:46

As multi-agent systems grow in importance for fields ranging from customer support to autonomous decision-making, OpenAI has introduced Swarm, an experimental framework that simplifies the process of building and managing these systems. Swarm, a lightweight Python library, is designed for educational purposes, stripping away complex abstractions to…

1
KV Cache Explained 4:19

8M ago4:19

4:19

In this episode, we dive into the intriguing mechanics behind why chat experiences with models like GPT often start slow but then rapidly pick up speed. The key? The KV cache. This essential but under-discussed component enables the seamless and snappy interactions we expect from modern AI systems. Harrison Chu breaks down how the KV cache works, h…

1
‘Job satisfaction will go up’: How generative AI is changing work 20:00

8M ago20:00

20:00

The technology will take on routine, dull work, says Alenka Grealish, principal analyst at Celent.

1
The Shrek Sampler: How Entropy-Based Sampling is Revolutionizing LLMs 3:31

8M ago3:31

3:31

In this byte-sized podcast, Harrison Chu, Director of Engineering at Arize, breaks down the Shrek Sampler. This innovative Entropy-Based Sampling technique--nicknamed the 'Shrek Sampler--is transforming LLMs. Harrison talks about how this method improves upon traditional sampling strategies by leveraging entropy and varentropy to produce more dynam…

1
Google's NotebookLM and the Future of AI-Generated Audio 43:28

8M ago43:28

43:28

This week, Aman Khan and Harrison Chu explore NotebookLM’s unique features, including its ability to generate realistic-sounding podcast episodes from text (but this podcast is very real!). They dive into some technical underpinnings of the product, specifically the SoundStorm model used for generating high-quality audio, and how it leverages a hie…

1
How banks’ use of generative AI has evolved over the past year 19:53

9M ago19:53

19:53

Financial institutions have dramatically increased their investment and trust in large language models since 2023. Kartik Ramakrishnan, Capgemini’s deputy CEO of Financial Services and head of banking and capital markets, shares the results of a recent report that analyzed these changes.

1
Exploring OpenAI's o1-preview and o1-mini 42:02

9M ago42:02

42:02

OpenAI recently released its o1-preview, which they claim outperforms GPT-4o on a number of benchmarks. These models are designed to think more before answering and handle complex tasks better than their other models, especially science and math questions. We take a closer look at their latest crop of o1 models, and we also highlight some research …

1
Upstart’s CEO Dave Girouard explains brighter outlook for rest of 2024 22:49

9M ago22:49

22:49

Advances in the company’s AI-based lending models have made them better at predicting risk, which has led to growth, he says.

1
Breaking Down Reflection Tuning: Enhancing LLM Performance with Self-Learning 26:54

9M ago26:54

26:54

A recent announcement on X boasted a tuned model with pretty outstanding performance, and claimed these results were achieved through Reflection Tuning. However, people were unable to reproduce the results. We dive into some recent drama in the AI community as a jumping off point for a discussion about Reflection 70B. In 2023, there was a paper wri…

1
Composable Interventions for Language Models 42:35

9M ago42:35

42:35

This week, we're excited to be joined by Kyle O'Brien, Applied Scientist at Microsoft, to discuss his most recent paper, Composable Interventions for Language Models. Kyle and his team present a new framework, composable interventions, that allows for the study of multiple interventions applied sequentially to the same language model. The discussio…

1
‘These models will always hallucinate’: Seth Dobrin on LLMsDek 18:23

9M ago18:23

18:23

Dobrin, founder of advisory firm Qantm AI and former global chief AI officer at IBM, warns that popular generative AI models were trained on the whole of the internet and hallucinate at an unacceptable rate.

1
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges 39:05

10M ago39:05

39:05

This week’s paper presents a comprehensive study of the performance of various LLMs acting as judges. The researchers leverage TriviaQA as a benchmark for assessing objective knowledge reasoning of LLMs and evaluate them alongside human annotations which they find to have a high inter-annotator agreement. The study includes nine judge models and ni…

1
‘Fraud is pervasive throughout the entire industry’: Crypto insider 21:34

10M ago21:34

21:34

Jake Donoghue, author of the book Crypto Confidential, shares some of the worst practices he saw as a founder of a cryptocurrency company.

1
Breaking Down Meta's Llama 3 Herd of Models 44:40

11M ago44:40

44:40

Meta just released Llama 3.1 405B–according to them, it’s “the first openly available model that rivals the top AI models when it comes to state-of-the-art capabilities in general knowledge, steerability, math, tool use, and multilingual translation.” Will the latest Llama herd ignite new applications and modeling paradigms like synthetic data gene…

1
Some banks are making a Faustian bargain with fintechs: Karen Petrou 18:57

11M ago18:57

18:57

Karen Petrou, the managing partner at Federal Financial Analytics and a long-time observer of banking and regulation, says banks need to do far more due diligence on potential fintech partners and exert more control over these relationships.

1
DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines 33:57

11M ago33:57

33:57

Chaining language model (LM) calls as composable modules is fueling a new way of programming, but ensuring LMs adhere to important constraints requires heuristic “prompt engineering.” The paper this week introduces LM Assertions, a programming construct for expressing computational constraints that LMs should satisfy. The researchers integrated the…

1
What military members need from their banks 25:05

11M ago25:05

25:05

Two veterans and executives at Armed Forces Bank – Tom McLean and Jodi Vickery – share the challenges they see their customers face and new products the bank has rolled out this year to better serve them.

1
Regulators are wise to be more careful’ after Chevron ruling 47:07

12M ago47:07

47:07

Gene Scalia, the banking lobby’s lawyer on retainer for a potential challenge to Washington’s capital reform effort, discusses the state of administrative law after the overturning of a key legal precedent.