01Platforms
Amazon Bedrock AgentCore
Act Now - What
- Amazon Bedrock AgentCore: Natera says its Amazon Bedrock AgentCore voice agent books mobile phlebotomy appointments with 100% tool-calling accuracy and sub-seven-second latency
- Why
- Natera's result depends on a specific voice architecture and progressive-trust authentication path.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Amazon Bedrock AgentCore uses AWS ML Blog: Natera says its Amazon Bedrock AgentCore voice agent books mobile phlebotomy appointments with 100% tool-calling accuracy and
Sources: AWS ML Blog
02Platforms
Amazon EC2
Act Now - What
- Amazon EC2: AWS reports that NVIDIA CUDA Multi-Process Service with Triton on Amazon EC2 reduced GPU infrastructure for automatic speech recognition by 75% while
- Why
- The reported capacity reduction depends on EC2, CUDA MPS, and Triton running together.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Amazon EC2 uses AWS ML Blog: AWS reports that NVIDIA CUDA Multi-Process Service with Triton on Amazon EC2 reduced GPU infrastructure for automatic speech recognition by
Sources: AWS ML Blog
03Platforms
Amazon Quick
Act Now - What
- Amazon Quick: GoDaddy says its two-year migration from a legacy BI tool to Amazon Quick saves 15,000 hours annually, cut its dashboard count by 50%, reduced rendering
- Why
- GoDaddy reports operational savings from a completed migration, which need a comparable local baseline.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Amazon Quick uses AWS ML Blog: GoDaddy says its two-year migration from a legacy BI tool to Amazon Quick saves 15,000 hours annually, cut its dashboard count by 50%
Sources: AWS ML Blog
04Platforms
Amazon SageMaker AI
Act Now - What
- Amazon SageMaker AI: Amazon SageMaker Python SDK v3 unifies script-mode workflows through ModelTrainer and ModelBuilder, while SourceCode can synchronize local code into any
- Why
- The SDK change targets iteration time between local code and runtime containers.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Amazon SageMaker AI uses AWS ML Blog: Amazon SageMaker Python SDK v3 unifies script-mode workflows through ModelTrainer and ModelBuilder, while SourceCode can synchronize
Sources: AWS ML Blog
05Platforms
AWS Agent Registry
Act Now - What
- AWS Agent Registry: AWS published Agentic Resource Discovery (ARD), an open specification for agent discovery, alongside AWS Agent Registry - a centralized, searchable catalog
- Why
- The registry is meant to make agent capabilities searchable and governed across environments.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. AWS Agent Registry uses AWS ML Blog: AWS published Agentic Resource Discovery (ARD), an open specification for agent discovery, alongside AWS Agent Registry - a
Sources: AWS ML Blog
06Platforms
Databricks AI Runtime
Act Now - What
- Databricks AI Runtime: Databricks frames goodput:the share of time large-scale training performs useful work:as the metric determining training efficiency
- Why
- The product claim centers on useful training time during fast, fault-tolerant execution.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Databricks AI Runtime uses Databricks AI: Databricks frames goodput:the share of time large-scale training performs useful work:as the metric determining training
Sources: Databricks AI
- What
- Docker: Amazon SageMaker Python SDK v3 unifies script-mode workflows through ModelTrainer and ModelBuilder, while SourceCode can synchronize local code into any
- Why
- The cited workflow avoids rebuilding images, which changes iteration speed only if runtime parity holds.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Docker uses AWS ML Blog: Amazon SageMaker Python SDK v3 unifies script-mode workflows through ModelTrainer and ModelBuilder, while SourceCode can synchronize local code
Sources: AWS ML Blog
- What
- Dependabot: GitHub says the GitHub Copilot app can automate repetitive Dependabot pull-request triage
- Why
- The cited GitHub workflow automates pull-request triage, where review quality matters as much as speed.
No material change · Evidence held steady into Act Now on 2 signals across 2 sources. Dependabot uses GitHub AI Blog: GitHub says the GitHub Copilot app can automate repetitive Dependabot pull-request triage
Sources: GitHub AI Blog · GitHub AI Blog
09Tools
GitHub Copilot
Act Now - What
- GitHub Copilot: GitHub says the GitHub Copilot app can automate repetitive Dependabot pull-request triage
- Why
- GitHub ties Copilot to a bounded maintenance task rather than unrestricted repository changes.
No material change · Evidence held steady into Act Now on 2 signals across 2 sources. GitHub Copilot uses GitHub AI Blog: GitHub says the GitHub Copilot app can automate repetitive Dependabot pull-request triage
Sources: GitHub AI Blog · GitHub AI Blog
10Tools
Agent Arena
Act Now - What
- Agent Arena: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at $0.24 and GPT 5.6 Luna at $0.07 against Claude
- Why
- The board shows wide cost and success differences, so a single leaderboard rank hides the tradeoff.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Agent Arena uses Agent Arena | AI Agent Performance Leaderboard: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High)
Sources: Agent Arena | AI Agent Performance Leaderboard
- What
- Alyx: Arize says its Signal tool found two hidden retry loops in the production agent Alyx: a duplicate task-state loop and a 43-call dataset retry that had
- Why
- Alyx produced tool activity that looked valid while retries consumed work without progress.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Alyx uses Arize AI Blog: Arize says its Signal tool found two hidden retry loops in the production agent Alyx: a duplicate task-state loop and a 43-call dataset retry
Sources: Arize AI Blog
12Tools
Amazon Bedrock AgentCore Evaluations
Act Now - What
- Amazon Bedrock AgentCore Evaluations: uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with LangGraph, LlamaIndex, the OpenAI Agents SDK, Amazon Bedrock
- Why
- The cited release uses an OpenTelemetry contract across multiple agent frameworks.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Amazon Bedrock AgentCore Evaluations uses AWS ML Blog: uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with LangGraph
Sources: AWS ML Blog
13Tools
Claude Agent SDK
Act Now - What
- Claude Agent SDK: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with LangGraph, LlamaIndex
- Why
- AgentCore lists the SDK among supported frameworks, making cross-framework evaluation testable.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Claude Agent SDK uses AWS ML Blog: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with
Sources: AWS ML Blog
- What
- Genie One: Databricks announced Genie One features aimed at extending AI-assisted analytics from answering questions to taking action on insights
- Why
- Databricks is moving Genie from answering questions toward taking action on the resulting data.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Genie One uses Databricks AI: Databricks announced Genie One features aimed at extending AI-assisted analytics from answering questions to taking action on insights
Sources: Databricks AI
- What
- Google ADK: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with LangGraph, LlamaIndex
- Why
- The release names Google ADK as a supported framework for framework-neutral evaluation.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Google ADK uses AWS ML Blog: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with
Sources: AWS ML Blog
- What
- LangGraph: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with LangGraph, LlamaIndex
- Why
- The cited evaluation layer claims framework-neutral scoring for LangGraph agents.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. LangGraph uses AWS ML Blog: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with LangGraph
Sources: AWS ML Blog
- What
- LlamaIndex: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with LangGraph, LlamaIndex
- Why
- Shared criteria make a framework comparison possible without changing the business task.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. LlamaIndex uses AWS ML Blog: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with
Sources: AWS ML Blog
18Tools
ModelBuilder
Act Now - What
- ModelBuilder: Amazon SageMaker Python SDK v3 unifies script-mode workflows through ModelTrainer and ModelBuilder, while SourceCode can synchronize local code into any
- Why
- The new SDK pairs ModelBuilder with a unified script-mode workflow.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. ModelBuilder uses AWS ML Blog: Amazon SageMaker Python SDK v3 unifies script-mode workflows through ModelTrainer and ModelBuilder, while SourceCode can synchronize local
Sources: AWS ML Blog
19Tools
ModelTrainer
Act Now - What
- ModelTrainer: Amazon SageMaker Python SDK v3 unifies script-mode workflows through ModelTrainer and ModelBuilder, while SourceCode can synchronize local code into any
- Why
- ModelTrainer is part of SageMaker's attempt to unify previously separate script workflows.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. ModelTrainer uses AWS ML Blog: Amazon SageMaker Python SDK v3 unifies script-mode workflows through ModelTrainer and ModelBuilder, while SourceCode can synchronize local
Sources: AWS ML Blog
20Tools
OpenAI Agents SDK
Act Now - What
- OpenAI Agents SDK: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with LangGraph, LlamaIndex
- Why
- The cited AgentCore release includes the SDK in its framework-neutral evaluation path.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. OpenAI Agents SDK uses AWS ML Blog: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with
Sources: AWS ML Blog
- What
- Signal: Arize says its Signal tool found two hidden retry loops in the production agent Alyx: a duplicate task-state loop and a 43-call dataset retry that had
- Why
- Arize reports that Signal exposed retries previously disguised as valid tool activity.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Signal uses Arize AI Blog: Arize says its Signal tool found two hidden retry loops in the production agent Alyx: a duplicate task-state loop and a 43-call dataset retry
Sources: Arize AI Blog
- What
- SourceCode: Amazon SageMaker Python SDK v3 unifies script-mode workflows through ModelTrainer and ModelBuilder, while SourceCode can synchronize local code into any
- Why
- SourceCode promises faster iteration by changing runtime code without rebuilding the image.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. SourceCode uses AWS ML Blog: Amazon SageMaker Python SDK v3 unifies script-mode workflows through ModelTrainer and ModelBuilder, while SourceCode can synchronize local
Sources: AWS ML Blog
23Tools
Strands Agents
Act Now - What
- Strands Agents: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with LangGraph, LlamaIndex
- Why
- AgentCore names Strands among the frameworks supported by its shared scoring contract.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Strands Agents uses AWS ML Blog: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with
Sources: AWS ML Blog
24Tools
Triton Inference Server
Act Now - What
- Triton Inference Server: AWS reports that NVIDIA CUDA Multi-Process Service with Triton on Amazon EC2 reduced GPU infrastructure for automatic speech recognition by 75% while
- Why
- The cited AWS deployment couples Triton serving with GPU sharing to reach its capacity result.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Triton Inference Server uses AWS ML Blog: AWS reports that NVIDIA CUDA Multi-Process Service with Triton on Amazon EC2 reduced GPU infrastructure for automatic speech
Sources: AWS ML Blog
25Projects
OpenClaw
Act Now - What
- OpenClaw: GitHub describes OpenClaw as the fastest-growing project in GitHub history
- Why
- OpenClaw's rapid growth makes its early security work more useful than its adoption count alone.
No material change · Evidence held steady into Act Now on 2 signals across 2 sources. OpenClaw uses GitHub AI Blog: GitHub describes OpenClaw as the fastest-growing project in GitHub history
Sources: GitHub AI Blog · GitHub AI Blog
- What
- GitHub: published the lessons its team learned evaluating LLMs for real-world secret scanning, framing pre-production LLM evaluation as its own discipline that
- Why
- GitHub treats evaluation as a release gate for a real-world security workflow.
No material change · Evidence held steady into Act Now on 3 signals across 3 sources. GitHub uses GitHub AI Blog: published the lessons its team learned evaluating LLMs for real-world secret scanning, framing pre-production LLM evaluation as its own
Sources: GitHub AI Blog · GitHub AI Blog · GitHub AI Blog
- What
- OpenAI: TechCrunch reports OpenAI is building AI agents for everything and openly questions adoption : 'Will everyone use them?' : framing agent proliferation from
- Why
- OpenAI is increasing agent supply while the cited report leaves broad user demand unresolved.
No material change · Evidence held steady into Act Now on 3 signals across 3 sources. OpenAI uses AI News & Artificial Intelligence | TechCrunch: TechCrunch reports OpenAI is building AI agents for everything and openly questions adoption : 'Will everyone
Sources: AI News & Artificial Intelligence | TechCrunch · Arize AI Blog · Agent Arena | AI Agent Performance Leaderboard
28Models
Claude Opus 4.8
Act Now - What
- Claude Opus 4.8: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at $0.24 and GPT 5.6 Luna at $0.07 against Claude
- Why
- Agent Arena places cost and net task improvement on the same board, exposing route economics.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Claude Opus 4.8 uses Agent Arena | AI Agent Performance Leaderboard: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro
Sources: Agent Arena | AI Agent Performance Leaderboard
29Models
Claude Opus 5
Act Now - What
- Claude Opus 5: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at $0.24 and GPT 5.6 Luna at $0.07 against Claude
- Why
- Its listed task cost is materially higher than several alternatives on the same benchmark.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Claude Opus 5 uses Agent Arena | AI Agent Performance Leaderboard: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High)
Sources: Agent Arena | AI Agent Performance Leaderboard
- What
- Kimi K3: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at $0.24 and GPT 5.6 Luna at $0.07 against Claude
- Why
- Its Agent Arena placement must be read with both task success and cost, not rank alone.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Kimi K3 uses Agent Arena | AI Agent Performance Leaderboard: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at
Sources: Agent Arena | AI Agent Performance Leaderboard
31Companies
Amazon Web Services
Act Now - What
- Amazon Web Services: AWS reports that NVIDIA CUDA Multi-Process Service with Triton on Amazon EC2 reduced GPU infrastructure for automatic speech recognition by 75% while
- Why
- AWS reports lower infrastructure use on a specific MPS and Triton configuration, not a general cloud result.
No material change · Evidence held steady into Act Now on 2 signals across 2 sources. Amazon Web Services uses AWS ML Blog: AWS reports that NVIDIA CUDA Multi-Process Service with Triton on Amazon EC2 reduced GPU infrastructure for automatic speech
Sources: AWS ML Blog · AWS ML Blog
- What
- Arize: Stuart Sy of OpenAI, interviewed by Arize, argues better models do not fix every agent failure because the bottleneck has moved off the model and onto
- Why
- Arize's cited interview places the current bottleneck outside the model itself.
No material change · Evidence held steady into Act Now on 2 signals across 2 sources. Arize uses Arize AI Blog: Stuart Sy of OpenAI, interviewed by Arize, argues better models do not fix every agent failure because the bottleneck has moved off the model
Sources: Arize AI Blog · Arize AI Blog
33Companies
Databricks
Act Now - What
- Databricks: frames goodput:the share of time large-scale training performs useful work:as the metric determining training efficiency
- Why
- Databricks defines goodput as productive training time, changing which failure time teams should measure.
No material change · Evidence held steady into Act Now on 2 signals across 2 sources. Databricks uses Databricks AI: frames goodput:the share of time large-scale training performs useful work:as the metric determining training efficiency
Sources: Databricks AI · Databricks AI
- What
- PyTorch: Databricks frames goodput:the share of time large-scale training performs useful work:as the metric determining training efficiency
- Why
- Databricks ties its runtime claim to the share of training time that produces useful work.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. PyTorch uses Databricks AI: Databricks frames goodput:the share of time large-scale training performs useful work:as the metric determining training efficiency
Sources: Databricks AI
35Research
CorporateBench
Act Now - What
- CorporateBench: is a human-validated, multi-task question-answering benchmark for enterprise document collections whose evaluation corpora exceed 230,000 documents
- Why
- The benchmark targets document complexity and confidentiality absent from simple synthetic tests.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. CorporateBench uses arXiv AI + CL: is a human-validated, multi-task question-answering benchmark for enterprise document collections whose evaluation corpora exceed
Sources: arXiv AI + CL
36Research
HarnessLens
Act Now - What
- HarnessLens: targets two failure modes in agent-harness evolution: wasting verification rollouts on unrelated behaviors and allowing aggregate scores to hide specific
- Why
- The paper targets wasted verification and regressions concealed by aggregate scores.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. HarnessLens uses arXiv AI + CL: targets two failure modes in agent-harness evolution: wasting verification rollouts on unrelated behaviors and allowing aggregate scores
Sources: arXiv AI + CL
- What
- SARA: A paper on tool-augmented LLM agents argues that untrusted runtime observations can become action-driving commands and cause real-world side effects beyond
- Why
- The paper shows how untrusted runtime text can drive actions beyond the user's intent.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. SARA uses arXiv AI + CL: A paper on tool-augmented LLM agents argues that untrusted runtime observations can become action-driving commands and cause real-world side
Sources: arXiv AI + CL
38Technologies
Agentic Resource Discovery
Act Now - What
- Agentic Resource Discovery: AWS published Agentic Resource Discovery (ARD), an open specification for agent discovery, alongside AWS Agent Registry - a centralized, searchable catalog
- Why
- The specification addresses discovery across environments, where stale ownership can create unsafe routing.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Agentic Resource Discovery uses AWS ML Blog: AWS published Agentic Resource Discovery (ARD), an open specification for agent discovery, alongside AWS Agent Registry - a
Sources: AWS ML Blog
39Technologies
CUDA MPS
Act Now - What
- CUDA MPS: AWS reports that NVIDIA CUDA Multi-Process Service with Triton on Amazon EC2 reduced GPU infrastructure for automatic speech recognition by 75% while
- Why
- The AWS result attributes lower GPU use to MPS combined with Triton under a measured workload.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. CUDA MPS uses AWS ML Blog: AWS reports that NVIDIA CUDA Multi-Process Service with Triton on Amazon EC2 reduced GPU infrastructure for automatic speech recognition by 75%
Sources: AWS ML Blog
40Technologies
OpenTelemetry
Act Now - What
- OpenTelemetry: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with LangGraph, LlamaIndex
- Why
- OpenTelemetry is the contract that lets the cited evaluator compare agents across frameworks.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. OpenTelemetry uses AWS ML Blog: Amazon Bedrock AgentCore Evaluations uses OpenTelemetry as a framework-agnostic scoring contract and can evaluate agents built with
Sources: AWS ML Blog
- What
- Alice: Raises $140M at Near $1B Valuation as AI Security Revenue Nears $100M ARR - The SaaS Sentinel - Funding round: Series of funding totaling $140 million led
- Why
- Alice sells model testing and runtime monitoring, so detection quality must precede vendor adoption.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Alice uses saassentinel.com: Raises $140M at Near $1B Valuation as AI Security Revenue Nears $100M ARR - The SaaS Sentinel - Funding round: Series of funding totaling
Sources: saassentinel.com
42Companies
Anthropic
Act Now - What
- Anthropic: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at $0.24 and GPT 5.6 Luna at $0.07 against Claude
- Why
- The leaderboard evidence shows model cost and task success moving independently.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Anthropic uses Agent Arena | AI Agent Performance Leaderboard: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at
Sources: Agent Arena | AI Agent Performance Leaderboard
43Companies
Conviva
Act Now - What
- Conviva: CEO Keith Zubchevich said enterprises deploying more AI agents must evaluate the experience those agents create, not merely whether they work, because
- Why
- Conviva argues that functional completion misses the experience created by dynamic agent conversations.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Conviva uses SiliconANGLE theCUBE: CEO Keith Zubchevich said enterprises deploying more AI agents must evaluate the experience those agents create, not merely whether
Sources: SiliconANGLE theCUBE
44Companies
DeepSeek
Act Now - What
- DeepSeek: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at $0.24 and GPT 5.6 Luna at $0.07 against Claude
- Why
- Agent Arena shows a low-cost route with measurable improvement, but the result is benchmark-specific.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. DeepSeek uses Agent Arena | AI Agent Performance Leaderboard: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at
Sources: Agent Arena | AI Agent Performance Leaderboard
45Companies
GoDaddy
Act Now - What
- GoDaddy: says its two-year migration from a legacy BI tool to Amazon Quick saves 15,000 hours annually, cut its dashboard count by 50%, reduced rendering times to
- Why
- The cited savings combine fewer dashboards, faster rendering, and broader self-service access.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. GoDaddy uses AWS ML Blog: says its two-year migration from a legacy BI tool to Amazon Quick saves 15,000 hours annually, cut its dashboard count by 50%, reduced rendering
Sources: AWS ML Blog
46Companies
Google DeepMind
Act Now - What
- Google DeepMind: announced a pilot that it describes as the world's first double-blind AI evaluations
- Why
- The pilot tests whether brand knowledge changes evaluation of frontier-model outputs.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Google DeepMind uses Google DeepMind Blog: announced a pilot that it describes as the world's first double-blind AI evaluations
Sources: Google DeepMind Blog
47Companies
Instinct
Act Now - What
- Instinct: Viral AI startup Instinct has raised $350M at a $2.5B valuation | TechCrunch Summary: - Funding round/status: Series B funding; the latest round brings
- Why
- The funding evidence is strong while the product remains private and depends on broad app access.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Instinct uses techcrunch.com: Viral AI startup Instinct has raised $350M at a $2.5B valuation | TechCrunch Summary: - Funding round/status: Series B funding; the latest
Sources: techcrunch.com
48Companies
Keenable
Act Now - What
- Keenable: Startup Keenable Announces $26M in Funding To Build Web Search Infrastructure Built For AI Agents Keenable announced a seed funding round of $26 million
- Why
- Keenable is building machine-focused retrieval, which should be judged on agent answers rather than human browsing.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Keenable uses theaiinsider.tech: Startup Keenable Announces $26M in Funding To Build Web Search Infrastructure Built For AI Agents Keenable announced a seed funding round
Sources: theaiinsider.tech
49Companies
Moonshot
Act Now - What
- Moonshot: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at $0.24 and GPT 5.6 Luna at $0.07 against Claude
- Why
- Agent Arena's shared task board exposes whether lower cost survives an owned workload.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Moonshot uses Agent Arena | AI Agent Performance Leaderboard: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at
Sources: Agent Arena | AI Agent Performance Leaderboard
- What
- Natera: says its Amazon Bedrock AgentCore voice agent books mobile phlebotomy appointments with 100% tool-calling accuracy and sub-seven-second latency using a
- Why
- The result combines tool accuracy, response latency, and progressive authentication in one workflow.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Natera uses AWS ML Blog: says its Amazon Bedrock AgentCore voice agent books mobile phlebotomy appointments with 100% tool-calling accuracy and sub-seven-second latency
Sources: AWS ML Blog
- What
- NVIDIA: AWS reports that NVIDIA CUDA Multi-Process Service with Triton on Amazon EC2 reduced GPU infrastructure for automatic speech recognition by 75% while
- Why
- The reported infrastructure reduction is tied to a concrete serving configuration and latency target.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. NVIDIA uses AWS ML Blog: AWS reports that NVIDIA CUDA Multi-Process Service with Triton on Amazon EC2 reduced GPU infrastructure for automatic speech recognition by 75%
Sources: AWS ML Blog
- What
- Rillet: turned an unsolicited board update into a $1 billion accounting startup in 48 hours - Startup Fortune Summary: - Funding round and status: $100 million
- Why
- Rillet's funding follows adoption of AI-native accounting inside a core finance system.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Rillet uses startupfortune.com: turned an unsolicited board update into a $1 billion accounting startup in 48 hours - Startup Fortune Summary: - Funding round and status
Sources: startupfortune.com
53Companies
Tencent
Act Now - What
- Tencent: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at $0.24 and GPT 5.6 Luna at $0.07 against Claude
- Why
- A shared leaderboard result is useful only when the task mix matches the operating workload.
No material change · Evidence held steady into Act Now on 1 signal across 1 source. Tencent uses Agent Arena | AI Agent Performance Leaderboard: On the same Agent Arena board, P50 cost per task spans two orders of magnitude - DeepSeek V4 Pro (High) at
Sources: Agent Arena | AI Agent Performance Leaderboard