Topic

vendor evaluation

Every THE D[AI]LY BRIEF article on vendor evaluation — enterprise AI analysis, benchmarks, vendor comparisons, and ROI frameworks for technology and business leaders. Updated as new coverage publishes.

agent evaluation

Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.

The Era by Eon Benchmark ran nine models over 33 questions three times each and published the 95% intervals alongside the leaderboard. They are about 30 points wide, and 33 of the 36 pairwise comparisons could not be called. Your internal bake-off has the same problem.

September 10, 2026 · 13 min read
AI vendor shortlist

A Demo Vendor Is Perplexity's #3 Source. Rank the Domains.

Two studies published the same day measured what grounds AI software recommendations. Perplexity's third-most-cited domain is a demo-software vendor's marketing blog, 59.8% of its citations sit below Tranco rank 100,000, and 34.7% of numeric citations point at a page that won't open or doesn't contain the number.

September 2, 2026 · 13 min read
enterprise search

Glean vs Copilot vs Dust: Buy the Permission Model

Microsoft now bundles Copilot Search free with a $30 Copilot seat and reaches over 100 connectors, so Glean has to beat it on the remainder rather than in the abstract. Dust has 12 connectors and replaces source-system ACLs with its own workspace spaces.

August 28, 2026 · 17 min read
StartupBench

Agents Averaged 73. Only 30% Were Usable. Grade Pass/Fail.

StartupBench ran nine frontier agents across 97 workflows taken from products AI startups already sell. A 73.67 rubric average produced an acceptable deliverable on 29.55% of runs — and agents satisfied auxiliary requirements more often than core ones.

August 22, 2026 · 14 min read
AI agent observability

Best AI Agent Monitoring: Langfuse, Then a Real Kill Switch

Seven agent monitoring tools priced against the same workload: 50,000 runs a month at 12 steps each. Self-hosted Langfuse wins on cost and audit depth — but only one of these tools sits in the request path where it can actually stop a runaway agent, and that is the part you have to build yourself.

August 11, 2026 · 16 min read
voice agents

Eleven Voice Agents, One Bank Call, No Clean Winner

An arXiv benchmark of eleven production voice-agent stacks ran 3,300 simulated calls against a real database. The researchers found task completion collapses on complicated calls, per-call cost varies thirteenfold with no link to quality, and a verification leak hides inside correct refusals.

August 1, 2026 · 13 min read