Weekly AI Digest

Claude Opus 5.5, GPT-6, and Grok 4.7 All Landed This Week — and That Was Just Monday

Three frontier models, an agents API going GA, a €20B European merger, and EU AI Act enforcement arriving at last. Welcome to the densest AI week of 2026.

25 September 2026  ·  Boris Agatić  ·  AI Workshop Zagreb

September 22 will be remembered as the day the frontier model race fully lost its sense of orderly sequencing. Within a 48-hour window, Anthropic released Claude Opus 5.5 with near-superhuman software engineering scores, OpenAI launched two variants of GPT-6, and xAI followed Grok 4.7's September 21 debut with expanded API access. Meanwhile, Xiaomi's MiMo-V2.6-Pro quietly demonstrated that a trillion-parameter open-weight model can be trained for approximately $3 million. Outside the model releases, the business story of the week was the Cohere–Aleph Alpha merger creating a $20B European sovereign AI champion — and the EU AI Act notifying 30+ companies that formal information requests had arrived.

The Model Cascade: Three Flagships in 48 Hours

The releases came in rapid succession and each merits individual attention before the combined picture can be assessed.

Claude Opus 5.5
Anthropic · September 22

89.9% SWE-bench Pro, 66.4% Terminal-Bench. Anthropic's new flagship for agentic workflows and coding. Priced at $4/$20 per million tokens (input/output). Designed explicitly for long-horizon software engineering tasks.

GPT-6 Sol
OpenAI · September 22

Reasoning flagship with 1.05M context window. Priced at $2/$10 per million tokens. Positioned against Claude Opus 5.5 for enterprise reasoning and agentic tasks. Benchmark details pending independent validation.

GPT-6 Luna
OpenAI · September 22

Ultra-low-cost companion to Sol at $0.10/$0.50 per million tokens, same 1.05M context. Targets high-volume, cost-sensitive enterprise workloads where Sol's capabilities exceed requirements.

Grok 4.7
xAI · September 21

2.1 trillion parameters, 71.0% on DeepSWE benchmark. Priced at $2/$6 per million tokens. Represents xAI's strongest software engineering model to date, competitive with Opus 5.5 on specific coding metrics.

MiMo-V2.6-Pro
Xiaomi · Open-weight

1 trillion parameter mixture-of-experts, trained for approximately $3M. Best-performing Chinese open-weight model on current benchmarks. Signals that frontier-class open models are approaching commodity training costs.

The practical implication for enterprise AI buyers is a market structure change: the gap between frontier closed models and the best open-weight alternatives is narrowing meaningfully. MiMo-V2.6-Pro's $3M training cost does not mean enterprises can train their own frontier model at that price — but it does mean that fine-tuned derivatives of strong open-weight models are increasingly credible alternatives to API-priced frontier models for specific domain workloads.

Pricing signal for enterprise buyers: GPT-6 Luna at $0.10/$0.50 per million tokens represents a 95% cost reduction compared to GPT-6 Sol — within the same model family, same context window. The capability gap between Sol and Luna will determine whether Luna is a genuine enterprise option or primarily a high-volume consumer play. For Croatian enterprises currently on GPT-4o or Claude Sonnet pricing tiers, the Luna pricing point merits an evaluation cycle before Q4 infrastructure commitments.

Claude Opus 5.5 in Depth: What 89.9% SWE-Bench Pro Actually Means

SWE-bench Pro is the current gold-standard benchmark for agentic software engineering — it tests whether a model can autonomously resolve real GitHub issues in production codebases, not just write code snippets in controlled conditions. At 89.9%, Claude Opus 5.5 can autonomously resolve nearly nine out of ten issues that experienced software engineers would typically spend hours investigating.

The Terminal-Bench score of 66.4% measures performance on open-ended terminal tasks — the kind of command-line and system administration work that agentic AI must handle when deployed in DevOps or infrastructure automation contexts. The 66.4% figure represents a substantial improvement over previous models and brings agentic terminal operation into the range where supervised production deployment becomes defensible for a broadening class of tasks.

For enterprise software teams, Opus 5.5's benchmark position translates into a concrete deployment question: which categories of engineering work are now automatable with acceptable risk, and which still require human-in-the-loop oversight? The 89.9% SWE-bench figure suggests that routine bug fixes, dependency updates, test generation, and refactoring tasks across well-documented codebases are now strong candidates for autonomous agent deployment. Complex architectural decisions and security-critical changes are not.

Agentic AI

OpenAI Agents API reaches general availability for all developers

The OpenAI Agents API — which provides programmatic access to agent orchestration, memory management, tool use, and multi-agent coordination — reached general availability this week, opening access beyond the enterprise early-access program to all developers. The GA release coincides with GPT-6's launch, meaning developers can now build production agents on GPT-6 Sol and Luna from day one. For enterprise buyers, Agents API GA reduces the procurement friction for agentic workloads: the infrastructure is now stable, documented, and covered by OpenAI's standard SLAs.

Enterprise

Factory raises $200M for enterprise coding agents

Factory, which builds autonomous software engineering agents for enterprise codebases, closed a $200M funding round this week. The raise values Factory at approximately $2B and reflects institutional conviction that the market for autonomous software engineering — validated by Opus 5.5 and GPT-6's benchmark results — is large enough to support a dedicated enterprise vendor. For large Croatian software firms and technology consultancies, Factory's trajectory is a useful market signal: enterprise software engineering automation is transitioning from experimental to investable.

Infrastructure

Temporal Technologies raises $550M for agent orchestration infrastructure

Temporal Technologies, which provides durable workflow orchestration infrastructure used to coordinate multi-step agentic processes, raised $550M in a round valuing the company at $3.2B. Temporal's infrastructure is increasingly the substrate on which enterprise agentic systems are built — it handles state management, retry logic, and coordination between agents in ways that prevent the reliability failures that plagued early agentic deployments. The size of the raise reflects how central orchestration infrastructure has become to the agentic AI stack.

Europe's Sovereign AI Moment: The Cohere–Aleph Alpha Merger

The most significant European AI business story of September 2026 was the announcement that Cohere — the Canadian enterprise AI platform — and Aleph Alpha — Germany's best-capitalised sovereign AI company — are merging to create a combined entity valued at approximately $20 billion. The €600M contribution from Germany's Schwarz Gruppe (the operator of Lidl and Kaufland) is the largest single private investment in European AI to date and gives the merged entity a rare characteristic: a large, strategically aligned anchor investor with direct enterprise customer relationships across European retail and logistics.

The strategic logic is legible. Aleph Alpha has regulatory credibility, German institutional relationships, and EU AI Act compliance positioning that took years to build. Cohere has enterprise API infrastructure, multilingual model capabilities, and North American enterprise traction. Neither can compete alone with OpenAI and Anthropic at frontier capability. Together, they can target the segment of European enterprises — particularly in regulated sectors — that will not or cannot use US-hosted frontier models due to data residency requirements, and that require a vendor able to deploy on EU-sovereign cloud infrastructure.

For Croatian enterprises evaluating European alternatives to US frontier models, the merged Cohere-Aleph Alpha entity is the most serious option that will emerge from 2026. It is unlikely to match Claude Opus 5.5 or GPT-6 on raw capability benchmarks in 2026 — but for workloads where data sovereignty, EU jurisdiction, and GDPR compliance are primary requirements, it will be the vendor with the deepest compliance infrastructure and the most credible EU-market track record.

Finance

Nscale files S-1 targeting $14.6B IPO valuation

Nscale, the UK-based AI infrastructure provider that operates large-scale GPU clusters for AI training and inference, filed its S-1 this week targeting a $14.6B IPO valuation. The filing represents the first major AI infrastructure company to pursue a public listing in the post-GPT-6 market environment. Nscale's customer base includes several European frontier model projects and enterprise AI deployments. The S-1 provides the first detailed public window into the economics of European AI infrastructure at scale.

EU AI Act: The First Formal Requests Arrive

The EU AI Act's enforcement timeline accelerated materially this week with the delivery of formal information requests to more than 30 companies. The requests are not yet enforcement actions — they are the procedural step that precedes an enforcement action, requiring recipients to document their AI systems, describe their risk classification methodology, and demonstrate compliance with applicable obligations. But their arrival marks the transition from the preparatory phase of EU AI Act compliance to the accountability phase.

The sectors targeted in the first wave of requests are exactly those identified as high-risk in the Act's original text: hiring and recruitment AI systems, credit scoring and lending decision systems, and medical triage applications. Companies operating AI systems in any of these three sectors — including Croatian HR-tech firms, fintech platforms, and health-tech companies — should treat the first-wave requests as a direct signal about their own compliance timeline, even if they were not among the recipients.

The December 2 prohibited-practices deadline is the next hard compliance event. Prohibited practices include biometric categorisation systems, social scoring by public authorities, and real-time remote biometric identification in public spaces. Any Croatian technology company operating or selling AI systems that touch these categories needs to have completed its assessment of compliance exposure before that date.

For Croatian AI vendors selling into EU markets: The formal information requests this week are a preview of what scaled enforcement looks like. If you are selling an AI product into hiring, credit, or healthcare workflows in Germany, Austria, or the Netherlands — the most active enforcement markets — you should already have a compliance dossier, not a compliance project. If you don't, the window for proactive engagement with your customers about compliance documentation is closing. An enterprise customer who receives an EU AI Act information request will immediately ask their AI vendors for their compliance documentation. Having that documentation ready is now a sales and retention requirement, not an optional risk management exercise.

Security

OWASP: 88% of organisations deploying agents have had security incidents

OWASP published a report this week documenting that 88% of organisations that have deployed production AI agents have already experienced at least one security incident attributable to the agent's behaviour. The incidents range from prompt injection attacks to unintended data exfiltration, unauthorised API calls, and credential leakage through agentic tool use. For enterprises evaluating agentic AI deployment on Opus 5.5, GPT-6, or Grok 4.7, the OWASP finding establishes that agent security is not a theoretical concern — it is a production reality that requires deliberate architecture, not just a security review of the base model.

The Open-Weight Frontier: MiMo-V2.6-Pro Changes the Cost Calculus

Xiaomi's release of MiMo-V2.6-Pro — a 1-trillion-parameter mixture-of-experts model trained for approximately $3 million — is the most significant open-weight model development of September 2026. The $3M training cost figure is striking: it represents a 100x cost reduction compared to what similar-scale models cost to train twelve months ago, and it demonstrates that the combination of improved training efficiency, optimised MoE architectures, and lower GPU compute prices is compressing frontier-class training costs faster than most analysts projected.

MiMo-V2.6-Pro achieves the best benchmark scores of any Chinese open-weight model, and sits meaningfully above the previous open-weight frontier (Meta's Llama 4 family) on code generation and reasoning benchmarks. Its open-weight release means that any organisation with the infrastructure to run a 1T-parameter model can fine-tune it on proprietary data without paying per-token API costs.

The practical relevance for Croatian and European enterprises is primarily in the fine-tuning and on-premise deployment scenarios. A Croatian bank that cannot use US-hosted API models for regulatory reasons, but that needs strong code and document analysis capabilities, now has an open-weight option that — after appropriate fine-tuning on Croatian-language data and domain-specific documents — could meet its requirements without any cross-border data transfer.

Croatia Corner: Three Startups, One Big Week

Croatian AI startups had a strong funding week, with three companies closing rounds that collectively demonstrate the breadth of the local ecosystem.

Zagreb · Seed

Codeplain closes $3M seed round

Codeplain, the Zagreb-based AI coding assistant company, closed a $3M seed round this week. Codeplain's product targets software development teams with context-aware code generation that integrates directly with enterprise codebases and enforces team-specific coding standards. The raise is notable timing — landing in the same week as Opus 5.5 and GPT-6 — and raises the question of how specialised coding tooling companies differentiate against frontier model providers who are integrating agentic coding capabilities directly into their core products. Codeplain's bet is that enterprise-specific context and standards enforcement is a durable differentiator that generic frontier models cannot replicate.

Zagreb · Series A

Hypefy raises $7.2M for AI-powered social commerce

Hypefy, which uses AI to match brands with content creators and optimise social commerce campaigns, closed a $7.2M Series A. The company's platform uses multimodal AI to analyse creator content, predict campaign performance, and automate the matching and briefing process for influencer marketing campaigns. The raise positions Hypefy as one of the better-capitalised Croatian B2B SaaS companies of 2026, and its use of multimodal AI reflects the broader shift from text-only to vision-enabled enterprise AI applications.

Zagreb · Recognition

Mediqcode wins Infobip Shift 2026

Mediqcode, which applies AI to medical coding and clinical documentation workflows, won the Infobip Shift 2026 competition — one of the most visible startup recognition events in the CEE region. Mediqcode's platform automates the ICD-10 and DRG coding workflow that hospital billing teams currently perform manually, reducing coding time and improving reimbursement accuracy. The Shift win provides significant visibility to European healthcare system buyers and follows a regulatory period in which AI-assisted medical coding has gained provisional acceptance in several EU member states.

Taken together, the three raises reflect the maturing distribution of the Croatian AI startup ecosystem: developer tooling (Codeplain), creative and marketing AI (Hypefy), and regulated-sector vertical AI (Mediqcode) are all represented, suggesting that the ecosystem has moved beyond early-stage experimentation into domain-specific application development.

The Week in Numbers

89.9%
Claude Opus 5.5 on SWE-bench Pro
$20B
Cohere + Aleph Alpha combined valuation
88%
Agent-deploying orgs with security incidents (OWASP)
$3M
Cost to train MiMo-V2.6-Pro (1T params)

What to Watch in the Coming Weeks

Three frontier models in a week — which one should your enterprise actually use?

Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and Grok 4.7 each have different pricing, capabilities, and compliance profiles. Choosing the right model for your workload — and building an architecture that doesn't lock you into a single provider — requires more than reading the benchmarks. We help Croatian and European enterprises navigate model selection, agentic deployment, and EU AI Act compliance. Certified Claude partner.

Talk to us