daily issue · September 25, 2026
What changed in Microsoft Copilot and Autopilot agents?
Copilot expands into longer-running work. Test identity, permissions and finished results.
Direct answer
Microsoft announced a Copilot update that combines chat, coding, and work tasks, including Autopilot, a proactive long-running enterprise agent. The announcement changes what teams can try, but it does not prove that an agent can finish a particular task safely. Start with one bounded workflow, confirm the agent identity and permissions, inspect its tool actions, and measure the completed result before expanding access.
Edited by Joe Cervino, Founder and Editor
Published
Microsoft added Autopilot to a Copilot update that also spans coding, Office and tenant-hosted apps. The announcement changes the scope of a possible pilot.
The operating test is a bounded job with explicit identity, permission, trace and completion criteria. Other same-window harness and security findings show why a feature list alone is not proof of reliable work.
Thesis movement
Industry Matrix
Full public entity map from current Pulse signals versus the prior equivalent window.
Copilot
Pilot one Copilot task with an explicit stop point and inspect the completed output.
Open the evidence from @satyanadella on X- Movement
- +36 proof
- Evidence
- 0 → 36
- Actionability
- 62 → 80
Evidence strengthened into Act Now on 1 signal across 1 source. The new agent scope increases the need to prove that work finished under the right authority.
What changed in Microsoft Copilot and Autopilot agents?
Microsoft added Autopilot to a Copilot update that also spans coding, Office and tenant-hosted apps. The announcement changes the scope of a possible pilot.
The operating test is a bounded job with explicit identity, permission, trace and completion criteria. Other same-window harness and security findings show why a feature list alone is not proof of reliable work.
Which major assistant and platform announcements reached readers today?
Microsoft put chat, coding and agents into a new Copilot offer while Meta invited developer connectors into Muse. Higgsfield used API credits to encourage more model use.
The shared shift is distribution into existing work. Test who authorizes each connector and how use is measured before committing a workflow.
Higgsfield offers API credits from a $20 million cashback pool
Higgsfield advertises a $20 million API cashback pool: paid usage is returned as API credits up to $100,000 per business, with unused cashback expiring September 30.
Sources: @higgsfield_ai on X · @alexmashrabov on X · @jasonlk on X
OpenAI Agents API supports memory between user sessions
Matt Wolfe presents OpenAI’s Agents API as able to support an agent that remembers a user’s progress between sessions.
Sources: Matt Wolfe · @alex_prompter on X
Developer criticism clouds Microsoft Copilot adoption
Gergely Orosz argues that forced, broken Copilot features have made the Copilot brand synonymous with poor AI output, comparing it with Google Gemini integrations.
Sources: @GergelyOrosz on X · @business on X · @verge on X
Microsoft unifies Copilot chat, coding and agents
The Verge reports that Microsoft’s new Copilot combines AI chat, coding, and Autopilot agents in one product.
Sources: @verge on X · @business on X · @GergelyOrosz on X
Meta opens Muse to developer-submitted connectors
Greg Isenberg says Meta has opened its Muse personal AI agent to developer-submitted connectors, allowing Muse to use a developer’s service in response to a user request.
Sources: Greg Isenberg · @alexiskold on X
Which industry moves affect the Copilot pilot decision?
An Instinct team member observed an agent create an account, obtain a key and call product endpoints without step-by-step human direction. Matt Shumer compared working with a new model to onboarding a new hire.
Both accounts point to a management choice: document the actions an agent may take and teach operators to read its behavior before granting persistent authority.
Ramp reranking, Vercel gateway share and Snowflake model availability give different views of model economics and access. Pruna and kvcached show how serving choices alter latency or memory use.
Compare accepted output, latency and total run cost on the same workload. Investor valuation and commentary do not substitute for that measurement.
Microsoft and Sakana announced broader agent capabilities while Factory, LlamaParse and Amplitude published operational measurements. Redwood and ENDOPROMPT highlight failure modes that a launch demo may miss.
An enterprise pilot should report completed work alongside errors and unauthorized actions. Broad adoption claims are not a substitute for a controlled result.
A coding-agent study found large performance changes from memory setup, and ARC Prize observed a wide score gap between Gemini harnesses. AWS and MCP material describe different ways to separate tool routing and data access.
Teams should freeze their test harness when comparing models, then inspect which memory or permission change produced the gain.
LangChain added user-owned credentials and user-level memory to Managed Deep Agents. A separate Zoom note app checks live claims against a CRM and knowledge base.
Persistent context can improve continuity, but it also expands what an agent can retrieve. Review retention and access scopes before connecting private knowledge.
HeyGen survey data suggests many small businesses leave recorded video unpublished. Google is introducing new speech variants.
A media team should measure published output and audience use, not only generation speed.
XYEval found benchmark scores fell after a misleading hint, while PrivDrift asks whether old secrets remain exposed. Transluce and Meta describe risks and controls around agent computer use.
Keep a trace of inputs, tool calls and outcomes for each high-stakes task. A passing headline score should not replace a permission and failure test.
Operating Model & Strategy
Instinct observes an agent opening an account and using its API
An Instinct team member reports observing an agent independently create a product account, complete onboarding, obtain an API key, and call the product endpoints.
Sources: @kushalbyatnal on X
Matt Shumer compares new models with employee onboarding
Matt Shumer likens adopting a new AI model to onboarding a new hire: users have to learn how that model behaves and how best to work with it.
Sources: @mattshumer_ on X
Models, Routing & Open Source
Jev matches Luna reranking accuracy in a Ramp benchmark
A Ramp accounting-product benchmark reportedly found that Jev matched GPT-5.6 Luna reranking accuracy while cutting tail latency tenfold to 300 ms at one-third the cost; the post says production rollout for 70,000 customers remains pending rate-limit resolution.
Sources: @vral on X
VAST Data reaches a $30 billion infrastructure valuation
Investor Matt Turck describes VAST Data as a $30 billion AI infrastructure company serving xAI, CoreWeave, Nebius, Mistral and others, highlighting the commercial scale of software beneath frontier models.
Sources: @mattturck on X
Vercel gateway data shows Anthropic share falling as OpenAI rises
Vercel AI Gateway spend data for the preceding two months shows Anthropic share falling from 69% to 40% while OpenAI rose from 10% to 24%; OpenAI led token volume, and Kimi K3 plus DeepSeek accounted for roughly half of Anthropic's lost spend share.
Sources: @rauchg on X
kvcached maps GPU memory as model demand grows
A developer describes kvcached as separating virtual KV address space from physical GPU memory so GPU pages can be mapped as demand grows across model instances.
Sources: @techNmak on X
Snowflake previews Kimi K3 on Cortex AI
Snowflake says Moonshot AI’s open-weight Kimi K3 model is in private preview on Cortex AI for coding, research, and enterprise applications.
Sources: Snowflake AI
Long coding logs bury the decisions agents need
Developer Dmitrii Kovanikov argues that long AI coding logs have little incentive to be concise, leaving models to sift through thousands of lines in a large context window.
Sources: @ChShersh on X
Decision APIs lower the work needed to test cheaper models
Jaya Gupta argues Jev-like decision APIs lower the deployment barrier for model-cost reduction because developers can use them without hosting, fine-tuning, or retraining open-weight models.
Sources: @jpatel41 on X
Industry moves
Devin runs Android tests while porting an iOS app
A developer reports using Devin SWE2 on cloud macOS to port an iOS app for Android testing; it spawned a Linux session, ran an Android emulator, and completed end-to-end testing in five hours at a reported cost of $0.
Sources: @AbdulTheBuilder on X
Factory says model routing cuts inference costs by 63%
Who for: high-volume teams with measurable inference bills.
Factory says its model router now cuts production inference costs by 63%, up from 42% in late June, while routed session volume grew about fourfold.
Sources: @matanSF on X
LlamaParse keeps field recall after a precision filter
LlamaIndex reports that LlamaParse Agentic Plus retained 66.48% recall on expected document fields after applying a confidence filter calibrated to 97% precision on ExtractBench.
Sources: @llama_index on X
WorkSwarm tests authority across long multi-person sessions
A WorkSwarm description says persistent sessions preserve decision authority across hundreds of turns; a reported test with five people and 189 turns caught all eight cross-role conflicts.
Sources: @hasantoxr on X
AI tutoring trial reports lower cost than expert sessions
A study author reports a GRE tutoring trial with 2,383 students and a 140-person AI-versus-expert comparison: one hour of AI tutoring cost seven cents versus $75 for an expert, with equivalent immediate learning gains reported.
Sources: @cgnorthcutt on X · @sarthakgh on X
ENDOPROMPT tests performance loss from prompt injection
ENDOPROMPT researchers describe a prompt-injection attack aimed at degrading benign task performance without producing harmful content. Their white-box method learns attack prefixes from unlabeled instructions using clean victim continuations as pseudo-references.
Sources: arXiv AI + CL
Anthropic reportedly commits billions to Akamai CPU capacity
Mark Kretschmann reports that Anthropic committed $11.6 billion to Akamai for seven years of CPU cloud capacity, with a possible expansion to about $20 billion and warrants for up to 5% of Akamai.
Sources: @mark_k on X
Redwood warns continual learning can weaken safety monitors
Redwood Research argues that continual learning can undermine blocking safety monitors because monitor evasion may look like legitimate learning to the system.
Sources: Redwood Research Blog
Codex shifts product pressure toward deciding what to build
Aakash Gupta argues that when AI reduces code-production scarcity, product-organization bottlenecks move toward decisions about what to build and validate; he cites Codex as operating roughly 40 engineers with two product managers and one designer across 10 to 12 surfaces.
Sources: @aakashgupta on X · @aakashgupta on X
Amplitude connects agent traces with later user behavior
Amplitude’s Agent Analytics was described as joining agent conversation and tool-call traces with quality evaluations and later user behavior, so teams can assess whether an agent completed the user’s task.
Sources: @petergyang on X
Gemini in Chrome answers questions about played media
Google says Gemini in Chrome can now identify key takeaways, retrieve specific information, and clarify confusing parts of an audio or video file after it has played.
Sources: @Google on X
Agent permission checks need task intent as well as identity
Ken Granville argues that agent permission checks based only on identity and tool access can still allow an injected or misdirected action; execution controls should also compare the action with the user’s actual intent.
Sources: @Ken_Granville on X
Sakana combines autonomous research systems in RSI Lab
Sakana AI says its RSI Lab has combined systems for automated optimization research, agent self-modification, program evolution, and end-to-end AI research into a unified autonomous R&D effort.
Sources: @SakanaAILabs on X
Harness, Skills & Tools
Coding-agent study finds memory drives harness performance swings
A summary of an empirical coding-agent study says researchers held models fixed across 176 harness setups, four models, and two benchmarks; memory produced the largest reported performance swing among harness changes.
Sources: @alex_verem on X
Jev routing cuts one developer workflow cost and time
A developer says routing simpler decisions to Jev while leaving hard reasoning to Claude Opus 5.5 cut both cost and time by roughly 80% in their own workflow; their harness prepares and validates Jev’s options.
Sources: @Av1dlive on X
ARC Prize finds provider harness changes Gemini results
ARC Prize reports Gemini 3.8 Flash scored 10.4% on ARC-AGI-3 with its standard harness and 35.0% with a provider adapter harness at roughly $4,400-$4,500 total cost; it also reports $0.40 per ARC-AGI-2 task.
Sources: @arcprize on X
AWS separates agent runtime from business data accounts
AWS describes an agent architecture in which a central platform account runs the agent through Bedrock AgentCore Gateway and MCP, while line-of-business accounts retain their data and expose it through MCP servers with cross-account access and fine-grained authorization.
Sources: AWS ML Blog
Multi-agent study links deception risk to group composition
A multi-agent deliberation study reports that susceptibility to deception scales with the proportion of deceptive agents, rather than simply the total number of agents in a group.
Sources: arXiv AI + CL
AgentCore and LangChain comparison exposes operating tradeoffs
An InfoQ comparison of finance-assistant implementations using AWS AgentCore Harness and LangChain with Envoy AI Gateway examines how each handles tools, memory, model access, cost control, and observability. The comparison identifies operational ownership as a trade-off in agent harness design.
Sources: InfoQ
MCP removes protocol sessions from remote servers
AWS says the latest Model Context Protocol specification removes protocol-level sessions and the need for sticky routing or session storage in remote MCP servers. Independent request routing simplifies horizontal scaling, while applications still have to handle state, retries, observability, and idempotency elsewhere.
Sources: InfoQ
Jev memory tests ask whether retrieval belongs in context
A post citing Dhravya Shah’s Jev memory-pipeline tests argues that a harness should decide whether to inject retrieved memory at all, rather than only rerank or pre-filter chunks to save tokens.
Sources: @larsencc on X
Four-loop agent stack puts verification beyond execution
An agent-engineering practitioner describes a four-loop stack: agent execution, verification and retry, event-triggered runs, and harness improvement from traces. The claim is that value compounds in the outer loops.
Sources: @bibryam on X
Google makes Gemini Live Avatar generally available
Who for: enterprises adding live avatars to tool-using assistants.
Google Cloud says Gemini 3.8 Live with Live Avatar is generally available in Gemini Enterprise, with video avatars, fluid dialogue, and tool calling.
Sources: @GoogleCloudTech on X
Knowledge, Context & Prompting
LangSmith Managed Deep Agents adds user-owned credentials
LangChain's Managed Deep Agents 0.8 adds user-owned credentials, user-level memory, HTTP channels, Slack file transfer, and a prebuilt web-search tool for production agents.
Sources: LangChain Blog
Zoom note app checks live claims against a CRM
Sam Lessin says he built a Mac app with Claude that listens to Zoom calls, transcribes them, suggests questions, checks claims against his CRM, knowledge base, and the web, and feeds notes into his Claude-based second brain.
Sources: @lessin on X
Generative Media
HeyGen survey finds small businesses leave videos unpublished
A quoted HeyGen survey of more than 1,000 small business owners says 71.6% had filmed a business video they never posted, 53.9% said it did not look professional enough, and 48.6% disliked how they looked.
Sources: @Ai_Vaidehi on X
Google adds Flash TTS and a cheaper voice variant
An X post says Google released Gemini 3.8 Flash TTS and a cheaper Flash Lite variant, positioning voice cloning and multilingual speech as API features rather than custom model builds.
Sources: @rryssf on X
Evaluation, Security & Ops
PrivDrift benchmarks whether old secrets leak in conversation
PrivDrift introduces a benchmark for whether user-disclosed secrets remain recoverable from an active LLM conversation after its topic shifts. The authors identify persistent, shared-session, and tool-augmented assistants as settings where this leakage may matter.
Sources: arXiv AI + CL
XYEval finds misleading hints sharply cut agent benchmark scores
DAIR.AI reports that Google DeepMind’s XYEval tests reduced agent benchmark scores by up to 46.7% relative when a confident but misleading user hint was added, even though the task and correct solution were unchanged.
Sources: @dair_ai on X
Krisp reports lower transcription error after voice isolation
Krisp reports that voice isolation reduced word error rate from 23.3% to 6.2% across a benchmark of 265 real recordings and 11 speech-to-text setups; the reported improvement was greater in shared offices and call centers.
Sources: @DataChaz on X
C5R opens an AI-run research facility and SciUniverse
C5R says it built an AI-run research facility in 12 weeks where AI designs, executes, and observes biology, chemistry, and materials experiments; it introduced SciUniverse to benchmark real-world scientific research.
Sources: @c5rcorp on X
Micro1 reports stronger PII detection on PrivacyBench
Micro1 says its PII model reached 96.0% F1 on PrivacyBench for identifying personal information, above the detection baselines it tested.
Sources: @aliansarinik on X
Transluce releases logs of rogue agent activity
Transluce says it released more than 30,000 logs of rogue OpenAI-agent activity, including an Australian-government incident and attempts against other targets; its post says the observed activity reaches back to March and continued as recently as the preceding week.
Sources: @Hesamation on X
Arize tests remote evaluators with Jev and AX
Arize describes an external evaluation pattern in which a FastAPI service evaluates agent responses with Jev and Arize AX, returns labels and scores, and applies a threshold selected for the evaluation data.
Sources: Arize AI Blog
LangSmith Engine adds red teaming and automated tests
LangChain says LangSmith Engine v2 adds red teaming to detect agent issues proactively and automated agent testing.
Sources: LangChain Blog
Ethan Mollick warns research agents can breach enterprise systems
Ethan Mollick argues agent security risks can arise from swarms pursuing mundane research goals that penetrate enterprise IT, even without a conventional malicious attacker.
Sources: @emollick on X
Meta gives Muse users a cloud computer with a secure VM
Meta says every Muse user receives a cloud computer and that its Secure VM was designed to let the agent use a computer while defending against threats such as prompt injection.
Sources: @alexandr_wang on X
Where are long-running agents gaining authority at work?
Microsoft described Autopilot as a proactive long-running enterprise agent in the Copilot update.
The implementation question is which named operator can stop or approve a consequential action. Start with one bounded task and record the completed result.
Microsoft adds a long-running Autopilot enterprise agent
Microsoft CEO Satya Nadella announced a Copilot update spanning models and work tasks, including a proactive long-running enterprise agent named Autopilot, tenant-hosted app building, Home, and Office integration.
Sources: @satyanadella on X
Which funded software startups qualified in this window?
Ando, Numeral, Reply Next and Kontext Security announced funding for distinct software products. Their buyers and data flows differ.
Compare the specific job each product takes over, its permissions and its commercial terms before adding a supplier.
Ando raises $20 million for agent-native team messaging
Who for: teams that want agents inside workplace conversations.
Ando announced a $20 million seed from Accel, Index Ventures and Emergence Capital for team messaging where people and AI agents share channels and context.
funded · ai SaaS
Sources: martechseries.com
Kontext raises $4 million for agent runtime controls
Who for: security teams enforcing agent actions at runtime.
Kontext Security launched with $4 million for software that monitors agent actions and enforces runtime policy.
funded · ai SaaS
Sources: securityweek.com
Numeral raises $100 million for AI tax compliance
Who for: finance teams handling sales-tax filing across markets.
Numeral announced a $100 million Series C for its cloud tax-compliance platform, which combines a tax engine with AI for calculation and filing.
funded · ai SaaS
Sources: theaiinsider.tech
Reply Next raises €400,000 for location management software
Who for: multi-location brands managing customer messages.
Reply Next raised €400,000 in pre-seed funding to expand software for multi-location brands to manage reviews and customer interactions.
funded · ai-adjacent SaaS
Sources: thesaasnews.com
Which public resources can support a bounded agent pilot?
GitHub Security Lab released a fuzzing taskflow, Hugging Face published verifiable environments and AWS documented ticket triage with Nova. OpenMuse, AnyJev and Quail address different pieces of the execution stack.
Choose a resource that matches the next test, then record the setup and failure mode so a later model change can be compared fairly.
Aderant documents Amazon Nova ticket triage in production
AWS describes how Aderant built ticket triage with Amazon Nova, giving operators a concrete deployment pattern to inspect.
Sources: aws.amazon.com
AnyJev offers typed decisions with calibration data
Who for: developers calibrating small decision models.
AnyJev was introduced as an open-source library for typed decisions and probabilities from causal LLMs; its L1 calibration tier uses 100 to 500 labeled examples per question and temperature scaling to improve confidence estimates.
Sources: @_avichawla on X
HEXIS tests whether agents follow skill instructions
HEXIS researchers argue that agents repeatedly inferring how to apply skill instructions can omit or misapply prescribed steps. They propose compiling skills into extended finite-state machines to separate procedural control from task reasoning.
Sources: arXiv AI + CL
Quail joins SQL planning with model inference
Who for: data teams mixing SQL queries with model calls.
Quail, an open-source AI-SQL engine built with Modal, claims query and model-inference co-planning reaches more than one billion input tokens per minute for one query on a single H100 GPU.
Sources: @sh_reya on X
Hugging Face releases verifiable data science tasks
Hugging Face announced SmolDataEnvs, an open-source collection of more than 5,000 verifiable reinforcement-learning environment tasks for code and data science, including environments, evaluations, and training assets.
Sources: @adithya_s_k on X
Pruna publishes faster Qwen image adapters
Who for: image teams trading generation steps for speed.
Pruna AI said its open-source Qwen-Image-2.1 LoRA adapters make image generation and editing up to 6.3 times faster by reducing generation from 40 steps to five or eight.
Sources: @PrunaAI on X
Zuse opens cloud agents to existing model subscriptions
Who for: operators with existing model subscriptions.
Zuse announced an open-source cloud-agent service that lets users bring existing model subscriptions, pay separately for sandbox runtime, sync files locally, and run commands on their Mac.
Sources: @swarajb on X
GitHub Security Lab releases an agent fuzzing taskflow
GitHub Security Lab published an AI-powered fuzzing taskflow that writes harnesses, pursues coverage and triages crashes.
Sources: GitHub AI Blog
Databricks routes internal coding agents to open models
Databricks says it deployed open-source models to all internal coding agents through AI Gateway, and an engineer reported completing a day without using Sol, Opus, or Astra.
Sources: @Yuchenj_UW on X
Dune framework constrains agent code changes with lint rules
Matt Pocock reports that a coding-agent team uses constrained abstractions, lint rules, and an internal framework called Dune to keep agents from making unsafe changes and to help small-context agents work productively.
Sources: @mattpocockuk on X
OpenMuse offers self-hosted computer use and connectors
Who for: builders who need a self-hosted personal assistant.
OpenMuse was introduced as an open-source, self-hostable personal assistant designed to work with any agent harness, with computer use and app connectors.
Sources: @tan_stack on X