OpenAI runs GPT-5.6 Sol at 750 tokens/second on Cerebras chips. Google Gemini reaches 1 billion monthly active users. Anthropic targets a $2 trillion October IPO. White House convenes AI labs for safety talks. Palantir Q2 revenue up 93%.
The second week of August 2026 confirmed that the AI industry has entered a phase where speed and scale — not just capability — define competitive position. OpenAI previewed GPT-5.6 Sol Ultrafast on August 13, running the frontier model at 750 output tokens per second (14× its standard speed) using Cerebras hardware, bringing real-time AI responsiveness to enterprise workflows that previously had to choose between intelligence and latency. One day earlier, Google's CEO Sundar Pichai announced that Gemini has crossed 1 billion monthly active users — its fastest product growth ever — while simultaneously releasing Gemini 3.7 Flash as a high-speed complement. Away from the model race, Anthropic confirmed it is targeting a $2 trillion IPO valuation in October, backed by a revenue trajectory that has seen run-rate figures pass $30 billion. And in Washington, the Trump administration convened OpenAI, Anthropic, and Google for the first formal White House AI safety dialogue of the year.
TL;DR: OpenAI previews GPT-5.6 Sol Ultrafast (750 tokens/s, 14× speed, Cerebras-powered, invite-only from Aug 13). Gemini crosses 1 billion monthly active users; ChatGPT at 1 billion weekly. Google releases Gemini 3.7 Flash. Anthropic targets $2T IPO valuation in October; run-rate revenue projected at $100-120B by year-end. White House convenes AI labs for voluntary safety testing framework. Palantir Q2 revenue $1.94B (+93% YoY). Meta releases Muse Glimmer (30B, Apache 2.0, multimodal). Anthropic appoints first Chief Global Affairs Officer.
On August 13, OpenAI opened a limited preview of GPT-5.6 Sol Ultrafast — a new inference mode that runs the GPT-5.6 Sol model at up to 750 output tokens per second, making it approximately 14 times faster than standard processing. The key distinction from previous speed-oriented releases: Ultrafast does not switch to a smaller or less capable model. It delivers the full frontier model faster, powered by Cerebras hardware — the wafer-scale silicon Cerebras Systems has developed specifically for high-throughput inference.
The practical implication is significant. Tasks that previously required users to choose between speed (using a smaller, faster model) and intelligence (using the frontier model with higher latency) can now run at frontier quality at near-real-time speeds. The targeted applications include incident response, customer service and support, financial market analysis, and e-commerce — all workflows where response latency directly affects the quality of human-AI collaboration.
The Ultrafast announcement has a direct consequence for enterprise AI architects: latency is no longer an acceptable reason to downgrade model quality. Teams that chose GPT-4o-mini or GPT-5.5 Flash for interactive applications purely because frontier models were too slow now have a path to frontier quality without the speed penalty. The invite-only access model signals OpenAI is managing capacity carefully — organisations with latency-sensitive workflows should register for access now, as widening will follow demand. At the same pricing tier as standard GPT-5.6 Sol, this is a capacity and access question, not a cost question.
Cerebras as infrastructure: OpenAI's choice to use Cerebras hardware for Ultrafast — rather than Nvidia GPUs — is notable. Cerebras wafer-scale silicon is optimised for memory bandwidth, which is the primary bottleneck for large-model inference throughput. This suggests that the current generation of Nvidia-centric inference infrastructure has real ceilings on token generation speed that alternative silicon can meaningfully exceed. For enterprises evaluating on-premises AI deployments, the Cerebras architecture deserves a direct assessment alongside standard GPU-based options.
Google CEO Sundar Pichai announced on August 11 that the Gemini app has crossed 1 billion monthly active users, making it the company's fastest-growing product in its history. The milestone puts Gemini on roughly the same scale as ChatGPT — which crossed 1 billion monthly app users in May 2026 and reached 1 billion weekly users in July — signalling that the consumer AI assistant market has consolidated around two dominant platforms with genuinely mass-market reach.
Alongside the user milestone, Google released Gemini 3.7 Flash — a high-speed inference model designed to complement Gemini 3.7 Pro. Flash prioritises throughput and low latency for agentic workloads and developer-facing APIs, directly competing with OpenAI's GPT-5.6 Sol Ultrafast in the high-speed inference segment. The simultaneous release of a speed-tier model from both OpenAI and Google in the same week is not coincidental — both companies are converging on the insight that enterprise adoption of agentic AI is gated as much by inference speed as by model intelligence.
For enterprise decision-makers: The AI assistant market has effectively bifurcated into two mass-market consumer platforms (Gemini and ChatGPT) with 1 billion+ users each, and a developer/enterprise API layer where the same models compete on price, speed, and specialisation. Choosing an enterprise AI platform now requires a clear-eyed assessment of which layer your use case lives in — and whether consumer-grade safety and data handling is acceptable for your workflows.
Anthropic's investors are targeting a $2 trillion or higher valuation for the company's planned October IPO — a figure that would rank Anthropic among the most valuable technology companies ever to list publicly. The target is anchored in a revenue trajectory that has accelerated beyond almost all forecasts: from a $9 billion run rate at end of 2025 to approximately $30 billion in April 2026, with analysts projecting $100–120 billion in annualised revenue by year-end 2026. If that trajectory holds, Anthropic would end 2026 out-earning every public software company except Microsoft.
The revenue growth is driven by three compounding forces: enterprise API adoption (accelerated by Claude Opus 5's launch and the Ode With Anthropic JV), consumer Max plan subscriptions (now at Claude Opus 5 by default), and a reseller and certified-partner ecosystem that has dramatically extended Anthropic's commercial reach without proportional headcount growth. Notably, Anthropic has reportedly passed OpenAI in revenue on a run-rate basis — a reversal of the competitive position that held from ChatGPT's launch through 2025.
Ahead of the IPO, Anthropic appointed Tino Cuéllar — former California Supreme Court Justice and president of the Carnegie Endowment for International Peace — as its first Chief Global Affairs Officer. The hire signals Anthropic's intent to engage proactively with governments and regulators at the highest levels as its technology becomes economically critical infrastructure in multiple jurisdictions. For enterprise clients, this appointment suggests Anthropic will have significantly more sophisticated policy engagement in major markets over the coming 12–18 months.
The Trump administration convened representatives from OpenAI, Anthropic, and Google at the White House this month to discuss a new US framework for voluntary AI safety testing. The framework, which stems from a June executive order on AI cybersecurity, outlines an opt-in approach to safety reviews of frontier models — an approach that differs structurally from the mandatory disclosure requirements being debated in the EU AI Act implementation process.
The opt-in structure reflects the current US administration's preference for industry-led safety commitments over regulatory mandates. For enterprise AI buyers, this has a practical implication: US AI safety credentials in 2026–2027 are likely to come primarily from voluntary commitments and third-party audits, not from regulatory compliance certifications equivalent to EU AI Act conformance assessments. Organisations operating in both US and EU markets will need to manage two different compliance frameworks in parallel.
OpenAI expanded its Daybreak cybersecurity programme this week, introducing GPT-5.6-Cyber — a specialised model for authorised security work — alongside Daybreak Blue (for defenders) and Daybreak Red (for offensive security research under controlled conditions). The initiative gives approved security professionals access to frontier models for vulnerability research, code review, incident response, and security testing, with enhanced safety guardrails for the security context.
Palantir Technologies reported Q2 2026 revenue of $1.94 billion, up 93% year-over-year — crushing analyst expectations and providing the most concrete publicly available evidence that enterprise AI adoption has moved from pilots to production at scale. US commercial revenue grew 149% year-over-year. Adjusted EPS came in at $0.41 against an expected $0.34, and Palantir raised its full-year 2026 guidance to $8.15 billion.
Palantir's results are a useful leading indicator for the broader enterprise AI market because the company's revenue is almost entirely derived from deploying AI in production — not from selling SaaS or consulting advisory. When Palantir grows at 93%, it reflects actual compute cycles running AI workloads in real enterprise and government environments. The acceleration in US commercial revenue (+149%) is particularly significant, suggesting that the largest US enterprises have moved decisively past the evaluation phase.
What Palantir's numbers mean for the rest of the market: If Palantir — which works primarily with the most analytically sophisticated organisations in defence, intelligence, healthcare, and financial services — is growing at 93% on production AI, the implication is that AI deployment in mainstream enterprises (which typically lag these sectors by 12–18 months) is still in early innings. The adoption curve for most businesses is not slowing; it is about to steepen.
750 output tokens/second (14× standard speed) using Cerebras chips. Frontier-quality model with near-real-time inference. Limited preview, invite-only. Same pricing as standard Sol.
High-throughput inference model for agentic workloads. Designed for speed-sensitive developer and enterprise API use. Complements Gemini 3.7 Pro. Available via Vertex AI and AI Studio.
30B-parameter dense multimodal model, Apache 2.0. Tuned for local agentic tool use, coding, and LLM-as-judge tasks. 131K context, supports 100+ languages. Strong on-premise deployment option.
Speed-optimised model from ByteDance's Seed family. Targeting developer and enterprise API markets with competitive pricing. Strong on Chinese-language and multilingual benchmarks.
Meta's Muse Glimmer deserves particular attention among open-weights releases this week. At 30 billion parameters under Apache 2.0 licensing, it is fully deployable on-premise — including on consumer-grade servers with sufficient RAM — and is specifically tuned for local agentic tool use and LLM-as-judge tasks. For organisations that cannot send data to cloud inference APIs due to data sovereignty or compliance requirements, Muse Glimmer represents one of the strongest open-weights options for local agentic deployment available today. Its 131K context window and 100+ language support make it viable for European multilingual enterprise use cases.
Anthropic confirmed an expanded strategic partnership with Google and Broadcom for multiple gigawatts of next-generation compute capacity. The partnership is structured to secure long-term inference infrastructure ahead of anticipated demand growth as Claude Opus 5 adoption accelerates. For enterprise clients, this signals Anthropic has secured the physical compute to support the API scale-up that a $2T company needs — reducing the risk of capacity-constrained API availability that has affected some providers in prior growth phases.
OpenAI appointed Dali Rajic as Chief Revenue Officer on August 13 — a signal that the company is building out the enterprise GTM infrastructure required to compete with Anthropic's accelerating revenue growth. The CRO role at OpenAI is newly created at the C-suite level, reflecting the shift from developer-led adoption (where product quality drives growth without a traditional sales motion) to enterprise-led adoption (where procurement, legal, and compliance require dedicated revenue infrastructure).
xAI's Grok Voice Think Fast 2.0 auto-migrated on August 5, moving from $0.05/min to $0.08/min pricing. The model improves latency and conversational fluency for voice-first applications — a segment where xAI has positioned Grok as a differentiator relative to text-primary ChatGPT and Claude. Users who preferred v1.0 pricing were given a window to pin the earlier version before migration.
Croatia's AI startup ecosystem continues to attract international attention, with Zagreb-based companies increasingly finding product-market fit in niche enterprise AI applications that larger platform vendors have not yet commoditised. Mindsmiths remains one of the most active Croatian AI companies in the enterprise space, developing tools for AI interaction design and decision intelligence — areas experiencing strong demand from financial services and telecommunications clients across the Adriatic region.
The broader investment picture remains solid: $125 million raised across six equity funding rounds through May 2026, with AI-adjacent companies in logistics, deep tech, and enterprise software among the primary recipients. What has changed materially in the past 60 days is the nature of inbound inquiries: Croatian companies that previously received interest primarily from regional investors are now fielding conversations with pan-European and US-based investors who are actively building out Central European AI portfolios ahead of anticipated EU AI Act compliance demand.
The EU AI Act is increasingly functioning as a market-creation mechanism for Croatian SMEs. Companies that have built compliance tooling, conformity assessment capabilities, or audit workflows on top of Claude or similar frontier models are finding that demand from Austrian, German, and Slovenian enterprises — which face the Act's requirements and prefer working with nearby partners who understand the regulatory context — is outpacing their capacity to deliver. For Croatian founders operating in the AI-adjacent space, EU AI Act compliance consulting and tooling is one of the clearest near-term revenue opportunities in the market.
For Croatian businesses considering AI investment: The convergence of Gemini at 1 billion users, ChatGPT at 1 billion weekly users, and Ultrafast inference puts consumer-grade AI tools at a capability level that most enterprise internal tools cannot match. The practical question is no longer whether to adopt AI, but how to do so in a way that is compliant with EU AI Act requirements, preserves data sovereignty, and generates defensible competitive advantage rather than merely matching what your competitors are doing with the same tools.
Our team helps Croatian and European businesses evaluate, deploy, and manage frontier AI — from model selection and speed-tier architecture to EU AI Act and DMA compliance. Certified Claude partner, no vendor lock-in.
Talk to us