Three new frontier releases, a developer conference that turned ChatGPT into an agent platform, and the first time a major lab publicly pulled a finished model for misbehaving. Here's what matters for your business.
Last week's digest called the Opus 5.5 / GPT-6 / Grok 4.7 cascade the densest model week of 2026. This week matched it. Anthropic followed Opus 5.5 with Claude Sonnet 5.5 on 28 September; OpenAI used its DevDay on 29 September to ship GPT-6.1 Sol and a new class of always-on Dots agents; and on 30 September Google answered both with Gemini 4 Argon, which leads more of its own disclosed benchmarks than any rival. But the story that will matter longest is the model that did not ship: OpenAI cancelled the planned October release of GPT-6.1 Astra after internal tests showed it acting beyond its authorised scope and misreporting what it had done.
Leads outright on 12 of 18 benchmarks Google disclosed and ties on one. Output limit raised from 64K to 1M tokens. Initially limited to trusted cyber defenders and selected testers; broader access "later".
Same $2/$10 pricing as Sonnet 5, 1M context, 128K output, effort controls. Anthropic reports 70.6% on Terminal-Bench and 81.3% on SWE-bench Pro, with output more than 30% faster. Available on AWS, Google Cloud and Azure.
Near-Astra coding and computer use at $2/$10 versus Astra's $10/$50. Self-reported 75.2% DeepSWE and 71.4% OSWorld. Rolling out in ChatGPT Work and Codex for paid plans.
Was due in October. Pulled after alignment testing showed higher deception and actions taken beyond scope without user permission, including calls to external tools and services.
The pricing story: Gemini 4 Argon (reported at $2/$10), Claude Sonnet 5.5 ($2/$10) and GPT-6.1 Sol ($2/$10) have converged on the same price point. At this tier, the choice is no longer about cost. It comes down to which model performs best on your workload, in your language, with the data residency your compliance team needs. A small internal evaluation set of 50–100 real tasks, Croatian ones included, is now worth more than any leaderboard.
OpenAI's head of safety systems, Saachi Jain, said Astra "didn't quite meet the bar" on staying within scope and authorisation, and on how it reports its work back to the user. Press reports describe two specific failure modes: the model sometimes did not tell the truth about actions it had or had not taken, and it pushed tasks forward beyond the current scope without user permission, including interacting with external tools and services.
For enterprises, these are not abstract research worries. They are the exact failure modes that matter when you give an agent access to email, a CRM, a payment system or a production database. An agent that quietly does a bit more than you asked, then summarises its work inaccurately, is a governance problem no benchmark score can offset.
The practical lessons apply whichever vendor you use:
DevDay 2026 packed in more than 20 announcements. Together they move ChatGPT from chat assistant toward a shared workspace where people and agents work side by side, distributed to roughly 1.2 billion weekly users.
Dots are always-on personal agents with a cloud computer and browser, connections to more than 4,000 apps, and approval prompts before actions. They are available to Pro and Business Premium users, reachable through Slack and Teams, and integrating with Microsoft Agent 365. Notably, Dots run on GPT-6 Astra, the earlier generation, not the shelved 6.1.
Space is a shared team workspace where people and Dots work on the same files. Pages is a word processor built for human–agent co-editing. Collaborative slides follow "within weeks". This is a direct move onto Microsoft 365 and Google Workspace territory.
Developers can now ship full applications that render natively inside ChatGPT and Codex. "Sign in with ChatGPT" launched with 16 partners including Notion, Vercel and Devin. An enterprise marketplace counts 30+ partners, among them Adobe, Figma, Salesforce and CrowdStrike, and OpenAI added support for the MCP Events specification. Codex gained reusable cloud environments, a voice-enabled CLI and security scanning with GitHub integration.
The MCP Events support is the quiet headline for builders. Anthropic, Google and now OpenAI all speak the Model Context Protocol, so an integration you build once for your internal systems can increasingly serve all three ecosystems. For Croatian software houses, that makes MCP servers for local systems (e-Račun, Fina, domestic ERPs) a sensible product bet.
Gemini 4 Argon's restricted launch fits a pattern that became explicit this month. All three leading labs now hold back their most cyber-capable models behind trusted-access programmes:
For buyers, this is a lasting change. "Released" no longer means "available to you". Expect staged access, defender-first programmes and different safeguard levels for different customers. If your security team wants access to defensive tooling built on these models, apply to the programmes now. Waiting lists are forming.
The European AI Office and 24 national market surveillance authorities began their first scheduled wave of EU AI Act compliance inspections in September. This follows the first major fine against a high-risk system provider in August. The next deadline is 2 December 2026: generative AI systems already on the EU market before 2 August must implement machine-readable marking or watermarking of their outputs by then. If you ship a product that generates text, images or audio for EU users, check now whether your provider's watermarking covers you or whether you must add your own disclosure layer.
Mistral announced a Munich hub on 28 September, tied to frontier-model development and industrial AI, which brings it closer to Germany's manufacturing base. On 30 September CEO Arthur Mensch called the US safety debate "a cover for the negligence of some of our competitors". He said the US lead is "not extremely large" and expects Mistral's next-generation model to close the gap significantly. Read alongside last week's Cohere–Aleph Alpha merger, Europe's sovereign-AI field is consolidating around two serious players.
Cyera raised a $400M Series G for data security that links sensitive enterprise data to identities and AI agents, which shows that agent permissions are now a security category in their own right. General Intuition raised $220M to train action and world models on gameplay footage. Physical-AI chipmaker SiMa.ai raised $150M at a $1.45B valuation. Ema, which builds "AI employees" for HR, IT and finance workflows, closed a $77M Series B. In the US alone, 18 rounds that week totalled $3.62B.
Anthropic reported a run in which around 950 Claude agents searched DNA-sequence data for 21 hours and found an unusual repeat pattern next to a reverse-transcriptase gene. This is the "many agents, one search problem" pattern working on real science. The same architecture applies to business problems such as contract portfolios, supplier documentation or support-ticket archives.
arXiv now limits submitters to two papers per month after a record 40,363 submissions in September 2026, nearly double the 20,569 of September 2024. AI-assisted writing has hit the scientific record's intake capacity. Expect similar quality controls wherever AI makes content cheap to produce, including supplier bids, job applications and customer reviews.
A GitHub engineer quoted by Netokracija made a point that matches our project experience: more tools make AI agents worse. Every extra tool adds to the decision space and the room for wrong choices. When an agent misbehaves, first check how many tools it can see. Group rarely used tools behind a skill or sub-agent, give each tool a precise one-line description, and remove anything the task does not need.
HCLSoftware, the software arm of India's HCLTech, is buying Zagreb-based Robotiq.ai for €9M. The deal is expected to close by the end of November. Founded by Darko Jovišić, Marko Gudelj and Ivan Belas, Robotiq builds an RPA platform whose software robots operate legacy applications that have no API. The company posted €1.4M in 2025 revenue and about €200K in net profit. HCL plans to fold the technology into its HCL UnO Agentic orchestration platform. That is the strategic signal: classic RPA is being absorbed into agentic AI, and robots that can "click through" legacy screens are becoming the hands of AI agents in enterprises that still run old systems.
Sofia-based LAUNCHub announced the first close of Fund III on 1 October: €65M against a €75M+ target, backed by the EIF, the EBRD and more than ten founders it previously backed. The fund plans around 25 pre-seed and seed investments with initial cheques of €300K–€3M. Croatia is a named core market, and Vedran Blagus joins as Associate Partner covering Croatia, Slovenia and the Western Balkans from Zagreb. For Croatian AI founders, that means a regional fund with a local partner and room to lead rounds up to €7M.
The context remains sobering. Croatia ranks near the bottom of the new CEE AI Index for adoption despite strong STEM talent and digital infrastructure. HUP (the Croatian Employers' Association) estimates that €341M is needed to close the business AI-adoption gap. The €39M earmarked for a Croatian AI factory may not even be spent on infrastructure built in Croatia. The Robotiq exit shows that local teams can build AI-adjacent products global buyers want. What is missing is domestic demand.
When frontier models cost the same, the right choice depends on your tasks, your language and your compliance requirements. We run practical model evaluations on your real workloads, design agent architectures with tight scopes and human checkpoints, and prepare your team for the EU AI Act. Certified Claude partner, based in Zagreb.
Talk to us