OpenAI's O4 Pro resets reasoning benchmarks with a 10M-token context window. The EU AI Act issues its first enforcement fine. Perplexity raises $500M. Anthropic ships Claude Sonnet 4.7. Mistral releases Large 3. Croatia's Tenderly crosses $40M ARR.
The final week of July 2026 delivered the densest set of frontier model releases since the January wave. OpenAI dropped O4 Pro on Tuesday — a reasoning-first model with a 10-million-token context window that overtook all public benchmarks within 24 hours of release. The EU used the same week to fire its first formal shot under the AI Act, issuing a €35 million fine against a major HR-tech platform for undisclosed automated hiring scores. Perplexity closed a $500 million round at an $18 billion valuation and announced a pivot toward enterprise search. Anthropic shipped Claude Sonnet 4.7, a meaningfully faster mid-tier update, and Mistral released Large 3 as a fully open-weights model under the Apache 2.0 licence. In Croatia, Tenderly — the Web3 dev platform that has been quietly building an AI simulation layer — crossed $40 million ARR.
TL;DR: OpenAI O4 Pro launches July 29 (10M-token context, new reasoning architecture, top benchmark scores). EU AI Act first enforcement fine: €35M against HR-tech vendor for undisclosed automated hiring decisions. Perplexity raises $500M at $18B valuation, pivots to enterprise search. Anthropic ships Claude Sonnet 4.7 — 40% faster than Sonnet 4.6, same capability tier. Mistral releases Large 3 under Apache 2.0. Cohere launches Command R+ 2.1 with agentic tool improvements. Google Cloud announces Gemini 3.5 Pro GA (general availability) pricing tiers. Croatia's Tenderly crosses $40M ARR with new AI simulation layer.
OpenAI's O4 Pro, released on July 29, is not just an incremental update to the O-series — it is a wholesale rethinking of how a frontier reasoning model handles long-context tasks. The 10-million-token context window (roughly 7.5 million words) makes it possible to load an entire mid-size corporate codebase, a year of financial filings, or a full clinical trial dataset into a single prompt without chunking or retrieval.
The architecture relies on what OpenAI is calling Hierarchical Reasoning Chains — an internal mechanism that segments long reasoning problems into nested subproblems, solves them independently, and assembles the result rather than maintaining a flat chain-of-thought across the full token budget. According to OpenAI's technical summary, this allows O4 Pro to maintain coherence over much longer context windows without the accuracy degradation typically seen past the 1M-token range.
O4 Pro scored 94.3% on GPQA Diamond (PhD-level science questions), a new public record, and 88.1% on FrontierMath, narrowing the gap to human expert performance significantly. On SWE-Bench Verified (software engineering tasks), it reached 71.4% — nearly 5 percentage points above the previous record held by the Gemini 3.5 Pro variant that Google released last week with an agentic scaffolding layer. On long-context recall benchmarks (NIAH variants over 1M tokens), O4 Pro was the first model to maintain above 95% accuracy without degradation.
O4 Pro is available in the ChatGPT Pro tier immediately for conversation use and via the API at $15 per million input tokens, $60 per million output tokens. A "cached" input rate of $3.75 per million applies to prompts where the prefix is unchanged across requests — important for enterprise workloads that load the same large system prompt repeatedly. O4 Pro Mini, a lower-cost distillation, is expected to follow in four to six weeks.
What it means for enterprise buyers: The 10M-token context eliminates the retrieval layer entirely for many document-heavy workflows — contract review, compliance auditing, and technical due diligence. The tradeoff is latency: O4 Pro is not a real-time model. First-token latency on maximum context runs is measured in tens of seconds. For batch processing use cases this is irrelevant; for interactive chat it is a material limitation.
The most consequential regulatory event of the week — and arguably the quarter — came not from a product launch but from the European Data Protection Board. On July 28, the EDPB issued a €35 million fine against a major HR-technology vendor (not named pending appeal proceedings, but widely reported in European press as a platform used by over 400 large employers across the EU) for violating Article 26 of the EU AI Act.
The violation: using an AI system that assigned automated risk scores to job applicants without disclosing to those applicants that such a system was in use, without providing meaningful human review of adverse decisions, and without registering the system in the EU database for high-risk AI applications as required under Annex III of the Act.
The €35 million figure represents 3.1% of the company's global annual turnover — well below the 6% maximum under Article 99(3) for prohibited AI practice violations, but well above the 3% floor for non-compliance with other obligations. More significant than the amount is the enforcement pattern it signals: the EDPB chose a systemic, high-volume, undisclosed automated decision-making system as its first target rather than a general-purpose AI model. This tells the market exactly where the enforcement risk sits in 2026 and 2027.
The EDPB fine puts four categories of AI use firmly in the crosshairs: (1) automated CV screening or scoring that affects hiring decisions, (2) performance monitoring systems that generate automated risk scores for employees, (3) customer credit or fraud scoring without disclosed logic, and (4) any system registered in Annex III categories that has not yet been logged in the EU AI Act database. If your organisation uses any third-party SaaS that performs these functions, verify that vendor's compliance status before September 2026 — the EDPB has indicated it is conducting sector-wide reviews of HR tech, financial services AI, and healthcare diagnostic AI simultaneously.
10M-token context, Hierarchical Reasoning Chains, 94.3% GPQA Diamond. Available in ChatGPT Pro and API. $15/$60 per 1M tokens. Optimised for long-context reasoning, not real-time interaction.
40% faster time-to-first-token than Sonnet 4.6, same capability tier, same pricing ($3/$15 per 1M tokens). Extended thinking now enabled by default for complex reasoning tasks. Improved tool-use reliability in agentic workflows.
Open weights under Apache 2.0 licence. 128K context. Leads among open models on most European-language benchmarks. Instruction-tuned and base variants available on Hugging Face. Runs in 2×A100 configuration.
Improved agentic tool-calling with multi-step planning, reduced hallucination rate on grounded retrieval tasks. Enterprise-focused. Cohere claims 60% reduction in tool-call errors versus Command R+ 2.0.
Anthropic's update to Sonnet 4.7 is less about capability ceiling and more about operational reliability. The 40% reduction in time-to-first-token is significant for customer-facing applications where perceived speed matters. The bigger change for enterprise users is that extended thinking is now on by default for prompts that trigger complex reasoning — previously, extended thinking required an explicit API flag. This means workloads that were silently underperforming on multi-step reasoning tasks may produce notably better outputs without any prompt changes.
Anthropic also published an updated system card noting a 22% reduction in refusal rate for legitimate professional queries in legal, medical, and financial contexts — a response to enterprise feedback that earlier Sonnet 4.x versions were too conservative for professional workflow automation.
Mistral's decision to release Large 3 under Apache 2.0 — fully open, commercially usable without restriction — is the clearest signal yet that the open-source frontier is catching up to closed APIs on European-language tasks. On multilingual benchmarks covering Croatian, Hungarian, Czech, Slovak, and Romanian, Mistral Large 3 outperforms every other open model and approaches the performance of GPT-5 class models on structured reasoning in those languages. For Central and Eastern European businesses considering on-premise AI deployment to address data residency requirements, this is the most compelling option yet available.
Perplexity AI closed a $500 million Series E round at an $18 billion valuation on July 26, led by SoftBank Vision Fund 3 with participation from Nvidia, Databricks, and several sovereign wealth funds. The round brings Perplexity's total raised to just over $1.2 billion since its 2022 founding.
More significant than the financing is the announced product direction: Perplexity is pivoting its primary growth focus from consumer search to enterprise knowledge retrieval. The new Perplexity Enterprise tier, launching in September 2026, will allow organisations to connect internal knowledge bases, Confluence wikis, SharePoint repositories, and Salesforce instances to Perplexity's retrieval engine — with answers cited, scoped to company data, and auditable.
Enterprise knowledge retrieval is one of the most consistently reported pain points in the AI adoption surveys run by McKinsey, Gartner, and Forrester over the past 18 months. The pattern is consistent: companies have deployed AI assistants, but employees cannot reliably find information locked in legacy systems, conflicting document versions, or unstructured email chains. Perplexity's approach — apply its citation-anchored search interface to internal data — is positioned directly against Microsoft Copilot for Microsoft 365 and Google's Gemini for Workspace on this specific use case.
After the dramatic launch of Gemini 3.5 Pro on July 17, Google Cloud announced this week the official general-availability pricing structure that enterprises have been waiting for before committing to production workloads. The pricing tiers: $7 per million input tokens, $21 per million output tokens for standard context (up to 128K); $14/$42 for the 1M–2M token range; cached inputs at $1.75. These rates position Gemini 3.5 Pro meaningfully below OpenAI's O4 Pro for comparable tasks, though the models target different use cases (Gemini 3.5 Pro for multimodal and deep reasoning; O4 Pro for maximum reasoning depth with very long context).
Google also announced that the Deep Think reasoning layer is now available in Vertex AI via a dedicated endpoint, allowing enterprise customers to invoke extended reasoning on a per-request basis without switching models entirely. This is particularly useful for legal and scientific workflows where most queries are standard but occasional complex queries require deep analysis.
A paper published this week by researchers at MIT and Stanford ("Distributed Speculative Decoding for Long-Context Transformer Inference", arXiv:2026.07891) describes a production-scale implementation of speculative decoding that achieves 3.1× inference speed improvement on models with context windows above 100K tokens, without any degradation in output quality. The technique uses a small "draft" model to propose likely token continuations in batches, which the larger model then verifies in parallel — reducing the number of sequential forward passes required.
What makes this paper notable is not the speculative decoding concept itself (which has been known since 2022) but the specific optimisations for very long contexts. At context lengths above 500K tokens, naive speculative decoding breaks down because the draft model's proposals diverge significantly from the large model's distribution. The paper introduces a context-adaptive drafting schedule that adjusts draft model confidence thresholds based on the accumulated context, maintaining the speedup even at multi-million-token lengths. OpenAI's technical blog noted this week that O4 Pro's inference infrastructure uses techniques "closely related" to the MIT/Stanford approach.
Microsoft confirmed O4 Pro will be available in Copilot Studio by mid-August 2026, with a dedicated "Document Mode" that loads up to 10M tokens of enterprise documents — contracts, policies, technical specs — for analysis workflows. Pricing will be consumption-based within existing Microsoft 365 E5 AI credits.
Meta published a comprehensive fine-tuning toolkit for Llama 4 Scout on Hugging Face, including QLoRA and full fine-tuning recipes, evaluation harnesses, and a set of domain-specific adapters (legal, medical, code) that serve as starting points. The toolkit dramatically lowers the barrier to domain-specific open-model deployment.
Nvidia's Q2 2026 earnings call confirmed that Blackwell Ultra GPU shipments to Microsoft, Google, Amazon, and Oracle are running approximately six weeks ahead of schedule, driven by improved yields at TSMC's N3B node. The supply acceleration is expected to lower GPU compute costs in cloud APIs through Q4 2026.
Cohere's Command R+ 2.1 is now the default model powering Salesforce Einstein for enterprise contract analysis and account research workflows. The integration gives Salesforce customers access to Cohere's grounded retrieval capabilities within existing Salesforce Data Cloud pipelines, without data leaving the Salesforce trust boundary.
Zagreb-based Tenderly, best known as the smart-contract simulation and debugging platform for Web3 developers, announced this week that it has crossed $40 million in annual recurring revenue — a milestone driven in large part by its AI Simulation Layer, launched in Q1 2026. The feature uses a fine-tuned code model to predict transaction outcomes, detect gas inefficiencies, and surface potential exploits in smart contracts before deployment, dramatically reducing the manual review burden on blockchain security engineers.
Tenderly's growth trajectory is notable for two reasons. First, it demonstrates that Croatian deep-tech companies can build developer infrastructure products at global scale from Zagreb — Tenderly now counts over 140,000 developer teams across 85 countries as customers. Second, the AI Simulation Layer represents a clear archetype for AI adoption in specialised technical domains: rather than replacing engineers, it amplifies their ability to catch problems early, compress review cycles, and handle routine analysis so senior engineers can focus on novel problems.
Tenderly CEO Pero Novak confirmed the company is not currently fundraising and plans to reach profitability before any future round. The company employs approximately 130 people, predominantly in Zagreb, with a small team in San Francisco.
For Croatian businesses: Tenderly's milestone is a reminder that the most durable AI-native businesses are often vertical platforms — deep domain knowledge, a specific workflow, and AI woven into the core product rather than bolted on. If you are building an AI strategy for 2027, the Tenderly model (own a workflow, automate the dull parts, amplify experts) is a more defensible architecture than a horizontal AI assistant layered on top of existing tools.
Our team helps Croatian and European businesses evaluate, deploy, and manage frontier AI — from model selection to EU AI Act compliance. Certified Claude partner, no vendor lock-in.
Talk to us