Colin Michaels

Loading latest stories

Article

This Week in AI: Astra and ChatGPT Images 2.5

GPT-6 Astra, ChatGPT Images 2.5, Meta Muse, OpenAI Agents API, AlphaGenome Atlas, Suno v6, and this week’s practical AI news explained.

By Colin Michaels - Sep 11, 2026

An original editorial illustration showing an AI workflow carrying a rough assignment through research and code tools to finished documents, spreadsheets, slides, and a website, with an approval gate before completion.

Trust & transparency

Evidence & disclosures

Not yet classified. This article has not yet been classified under the current editorial standard. It may predate the policy; do not assume a product was owned, tested, supplied, or independently verified unless the article says so.

Linked references: 16 explicit sources in the article.

How evidence labels and corrections work

This week, the useful AI race moved from impressive answers toward controlled finished work.

This issue covers Friday, September 4 through Thursday, September 10, 2026, in America/New_York.

Last week, GPT-6 Astra arrived with the sort of launch claims that make the rest of the internet briefly forget how calendars work. This week gave us the more useful part: people got access, started putting it into real workflows, ran into limits, and learned that “the smartest model” is still not the same thing as “the right model for every job.”

Then OpenAI released ChatGPT Images 2.5, which may be the more immediately useful upgrade for a lot of people. It is supposed to generate faster, preserve subjects more reliably, and make focused edits without casually redecorating the rest of your picture like an overconfident contractor.

The common thread is not raw intelligence. It is control. Astra is supposed to carry a complicated assignment further. Images 2.5 is supposed to change exactly what you asked it to change. Both are trying to close the gap between an impressive first answer and a finished result you can actually use.

Source note: I used Future Tools, TLDR AI, The Rundown AI, Ben's Bites, The Neuron, and ThursdAI as discovery and context sources, then checked release details against OpenAI, Microsoft, Meta, Google DeepMind, Suno, and Anthropic. Newsletter credit is for helping surface and frame the week—not for replacing verification. I did not run a controlled Astra or Images 2.5 benchmark for this issue.

TLDR

  • GPT-6 Astra launched on September 3, just outside this issue's window, but its wider rollout and first full week of use belong to this issue. Access still differs between Chat, Work, Codex, enterprise settings, and the API.
  • Astra's real promise is end-to-end work: research, computer use, coding, and polished documents, spreadsheets, presentations, and websites. OpenAI's benchmarks and customer quotes are useful starting points, not independent proof for your workflow.
  • In my own multi-prompt Little Nauts build, Astra felt faster and more consistent than GPT-5.6 across images, storylines, and games—but it consumed two full resets, possibly another, and still required purchased usage.
  • ChatGPT Images 2.5 rolled out September 8 with more precise editing, stronger subject preservation, Sketch, templates, image comments, prompt sharing, and what OpenAI says is up to 50% lower latency than Images 2.0.
  • Developers now get GPT-Image-2.5 Flare for faster everyday production and Sunburst for slower, higher-precision creative work.
  • OpenAI also launched a public-beta Agents API and a Data agent for ChatGPT Work, pushing its agent tools deeper into software and company data.
  • Meta launched Muse, a personal agent that can browse, send email, book travel, and make approved purchases from a dedicated cloud VM. Its privacy and safety protections are Meta's launch claims and need real-world scrutiny.
  • Google released AlphaGenome Atlas, Suno released its v6 music-model family with industry partners, and Microsoft put MAI-Image-2.6 and its Flash variant into public preview for developers.
  • Anthropic updated its account of four cyber-evaluation incidents, saying the models showed biased reasoning and recklessness after vague scope and an accidentally open internet path let them reach real systems.

Astra's Second Week Was More Important Than Its Launch Day

I covered the official Astra launch in the previous issue because it happened on September 3. I am not going to pretend it launched again just to give this post a bigger headline.

What changed during this coverage window was that Astra started reaching more people and products, and the gap between the launch page and a person's model picker became part of the story.

OpenAI describes Astra as its model for complex professional work. It combines reasoning, browsing, coding, computer use, and artifact creation, with a particular focus on carrying a multi-step job from the original request to a polished document, spreadsheet, presentation, analysis, website, or application.

That distinction matters. A chatbot can tell me how to build a blog package. A work model has to read the notes, research the claims, create the files, preserve the CMS contract, generate the visuals, and prove that the final folder still imports. The answer is not the product. The finished, checked artifact is the product.

An original editorial illustration showing an AI workflow carrying a rough assignment through research and code tools to finished documents, spreadsheets, slides, and a website, with an approval gate before completion.

OpenAI says Astra is better than GPT-5.6 Sol at following templates, pulling only relevant context into a result, handling changing instructions, and asking a focused question when missing information could change the outcome. It also publishes strong results for computer use, automation, coding, browsing, science, and cybersecurity.

Those are OpenAI's evaluations. Some tables include independent benchmark names, but the launch page still reflects the company's selected harnesses, settings, comparisons, and presentation. The useful conclusion is not “Astra wins everything.” It is “Astra deserves a test on the hard work you already do.”

Access Is Still More Complicated Than the Headline

The launch page said Astra would expand from a limited set of organizations to Plus, Pro, Business, and Enterprise users, plus the OpenAI API, Microsoft Azure, and AWS Bedrock. By the end of this issue's window, OpenAI's Help Center described a more specific product split: GPT-6 Pro in Chat was rolling out to Pro, Business, and Enterprise plans, while Plus users could get Astra in ChatGPT Work and Codex.

That is not necessarily a contradiction; products, plan labels, and phased rollouts can change on different schedules. It is a warning not to design a workflow from one launch sentence. Check the model picker in the product you actually use, your workspace administrator's settings, your region, and the current credit rules.

A rollout map showing GPT-6 Astra access branching across Chat, ChatGPT Work, Codex, the OpenAI API, Azure, and AWS Bedrock, with different plan and administrator gates.

For the API, OpenAI lists standard Astra pricing at $10 per million input tokens and $50 per million output tokens, with separate cache rates. Fast mode is listed at up to twice the speed for twice the standard price. A long-context request can also have different pricing rules, so the number that matters is still the cost of one completed job—including retries, tool calls, and the expensive hour spent confidently building the wrong thing.

The Astra Test I Actually Care About

I do not need another model to write a prettier paragraph about my project. I need it to stay oriented while the job gets messy.

My useful Astra test would look like this: give it a real source folder, a CMS schema, a visual brief, and a strict draft-only boundary. Then change one requirement halfway through. Ask a side question. Add a source. Reject one image. Make it rebuild the manifest. The model passes only if it keeps the original goal, incorporates the change, avoids overwriting the wrong files, and validates the final package.

A six-part finished-work scorecard for Astra: instruction fit, evidence, artifact quality, recovery after changes, full-task cost, and approval discipline.

That is where Astra could be a real step forward. It is also where a more expensive frontier model may lose to a smaller model with a better workflow. If a focused tool can finish the job faster and more predictably, the smarter choice is the boring one that works.

My First Week With Astra: Fast, Consistent, and Hungry

I did get to use Astra on a real creative build this week, even though I did not run a controlled benchmark. The project was Little Nauts, a children's corner of the Dreadnauts world with comics, storybooks, recurring characters, music, progress tracking, and interactive learning games.

This was not a one-prompt website. I used multiple prompts to build out the images, storylines, character continuity, and games over time. I still supplied the direction, made the creative decisions, checked the continuity, and approved what belonged in the project. Astra helped construct the work; it did not wake up one morning and invent Little Nauts by itself.

Compared with GPT-5.6, the biggest improvement I felt was consistency at speed. Astra was better at holding onto what the project was supposed to be while moving through different kinds of work. A storyline could lead into an image, then into an interactive feature, without the whole thing feeling as if a new assistant had walked into the room halfway through.

That is the good news. The less cheerful news is that Astra ate through my weekly usage at a startling rate. I used both of the full resets I had available. As I remember it, I may even have been given an extra reset, and I still ran out and had to purchase more usage.

That is personal experience, not a universal rate calculation. My project involved many prompts, images, code, checks, and revisions, so it was never going to be a light chat. I also cannot separate every bit of the burn into model reasoning, tool activity, long context, retries, or the way the product counted credits that week. But the practical result was very clear: Astra is a hungry one.

My advice is to save it for work that benefits from its strengths. Let Astra handle the complicated build, difficult bug, long research job, or messy multi-step assignment. Use a cheaper model for routine summaries, small rewrites, and jobs that do not need the frontier. The fastest model in the room is not saving you money if it drains the tank before lunch.

A field-note summary of GPT-6 Astra showing stronger speed, continuity, and complex-task execution balanced against rapid weekly usage burn and the need to route routine work elsewhere.

Other People Found the Same Tradeoff—and Some Found Different Ones

The most interesting week-one reaction was not a clean verdict. It was a pattern: Astra often felt more capable when people gave it a real job, but that capability could be expensive, literal, or uneven depending on the workflow.

Matt Wolfe at Future Tools said Astra changed his mind after a run of model launches that had felt incremental. His recommendation was still selective: keep a cheaper model as the everyday default and bring Astra in for long agent workflows, large codebases, computer use, or research where the difference can justify the cost.

Ben's Bites described frontier models such as Astra as better at understanding intent, which made fuzzier instructions more workable. Ben also said he was burning through limits quickly and was experimenting with using the frontier model as an orchestrator that briefs cheaper agents instead of doing every piece itself. That sounds remarkably close to my own conclusion: use the expensive brain where judgment and coordination matter, then route the routine work elsewhere.

The negative reports were not only complaints about price. In a large Reddit discussion about Astra's intelligence and intuition, some developers said it was excellent at executing a specific, well-defined request but weaker at choosing the right idea or inferring an unstated goal. Others reported the opposite. That disagreement matters because model behavior is shaped by the codebase, instructions, tools, reasoning setting, and the kind of collaboration a person expects.

Usage burn showed up repeatedly. The Neuron's weekend digest summarized developer Seth Rose's self-reported account of spending nearly 40 percent of a $100-plan token allowance in two hours. His later 48-hour comparison showed more calls and more tokens per day, with Astra taking a larger share of estimated cost than its share of tokens. Those figures come from one person's telemetry, not an OpenAI audit, but they make the same warning harder to dismiss as a quirk of my project.

Then there are the spectacular demos. One experiment had Astra autonomously complete the original *Portal* over roughly 24 hours. Tom's Hardware reported 3,336 tool calls and a headline token cost of $571.18, which the creator said was covered by a Codex Pro subscription. It is an impressive demonstration of persistence, and an equally impressive reminder to keep the meter visible.

My week-one verdict is therefore neither “Astra is overhyped” nor “Astra replaces everything.” It is much more useful: Astra can hold a complicated creative and technical project together better than the models I was using before, and it can burn through an allowance faster than I expected. Both things can be true at once.

Astra's Safety Story Is Part of the Product

OpenAI says Astra is its first broadly deployed model to reach the Critical cybersecurity capability level under its Preparedness Framework. In its safety overview, the company says the unsafeguarded model could find previously unknown vulnerabilities and develop exploits against protected systems. It also says Astra is more aligned overall than GPT-5.6 Sol, while noting that the model's written reasoning was harder to monitor in adversarial tests designed to make it conceal that reasoning.

That is an uncomfortable combination: more capable, more careful in many evaluations, and potentially harder to watch in some of the situations where monitoring matters most.

OpenAI says the production system uses stronger isolation, access controls, checkpoint security, prompt-injection defenses, reasoning-and-action monitoring, and automatic stops. Legitimate work may sometimes pause for review. In the API, a flagged task can stop instead of waiting for a person inside the conversation.

A layered safety system for advanced AI work showing model alignment, a hard sandbox, least-privilege access, live monitoring, an automatic stop, and a human approval gate.

I would rather have a legitimate run pause than discover that the model interpreted “clean this up” as permission to erase the evidence. But safety friction has to be measured too. A system that stops constantly can push people toward workarounds, weaker tools, or riskier settings. The honest question is not whether the guardrails ever interrupt work. It is whether they reduce dangerous behavior without making safe work impossible.

ChatGPT Images 2.5 Is Really an Editing Release

OpenAI introduced ChatGPT Images 2.5 on September 8. The company says it produces sharper details, more natural lighting, richer textures, better reference-subject preservation, and more reliable multi-turn edits. OpenAI also says generation latency can be up to 50% lower than Images 2.0.

The speed claim will get attention. The editing claim is the one I care about.

Image generation has been good at making a dramatic first draft for a while. It has been much less dependable when I ask for one small change and it responds by changing the face, moving the furniture, replacing the dog, and apparently relocating the house to a neighboring dimension.

An original editorial illustration showing the same reading room preserved across three stages while a green chair and then an orange lamp are added as targeted edits.

OpenAI says Images 2.5 is better at understanding what not to change. That can turn image generation from a novelty into a workflow. A creator should be able to preserve a person's appearance, a product's shape, a room's layout, or a brand composition while changing only the requested element.

That does not mean identity or product accuracy is solved. It means the failure rate may be lower. People, pets, medical details, real products, and anything used as evidence still deserve a visual comparison against the original.

Sketch, Comments, Templates, and the End of Prompt-Only Editing

In ChatGPT, Images 2.5 adds more direct creative controls:

  • Sketch lets a person draw a rough visual reference and describe the finished image.
  • Templates provide structured starting points for formats such as posters, merchandise, and product photography.
  • Image comments let a person point to a specific area and describe the requested change.
  • Prompt sharing lets someone share the idea behind an image so another person can make a version with their own details.

The important shift is that prompting is no longer the only control surface. Words are useful, but sometimes the clearest instruction is a rough box, a circle around the wrong object, or “put the chair right here.”

A four-step controlled image-editing loop: preserve the reference, select the target, make one change, and compare before approval.

This is also a healthier creative model than pretending the perfect prompt appears fully formed. Real design is iterative. You make something, notice the problem, adjust one part, and keep the good pieces intact.

Flare and Sunburst: Two Different Image Jobs

For developers, OpenAI released two API models.

GPT-Image-2.5 Flare is the default choice for most applications. OpenAI positions it for creator content, social media, product experiences, visual search, rapid prototyping, and high-volume generation. The company says Flare produces higher-quality images than GPT-Image-2 with 50% lower latency.

GPT-Image-2.5 Sunburst is the premium precision option for workflows where tighter control across edits matters more than raw speed. OpenAI points to production-ready campaigns and polished product imagery as examples.

A job-based comparison of GPT-Image-2.5 Flare for faster everyday production and Sunburst for slower, high-precision editing.

This is a sensible split. Drafting a dozen visual ideas and carefully revising the selected campaign image are different jobs. Paying premium-model prices for every throwaway draft is wasteful. Using the fastest option for a delicate identity-preserving edit can be just as wasteful if it forces three retries.

OpenAI says Images 2.5 is available across ChatGPT, ChatGPT Work, and Codex tiers on desktop, mobile, and web, with the two models available in the API. Templates were not yet available in Work mode according to the September 8 ChatGPT release notes, so feature parity still depends on the product surface.

Better Images Need Better Receipts

OpenAI says it continues to use C2PA metadata and invisible watermarking to help identify images made with its tools. That is useful, but no provenance system is a magic truth detector. Metadata can be removed, screenshots can flatten it, platforms can discard it, and a labeled synthetic image can still be used in a misleading context.

For this post, I kept the generated illustrations separate from the deterministic explainers, labeled both in their captions, and recorded the source files and generation purpose in the image manifest. None of these images is presented as a screenshot, a hands-on benchmark, or evidence that a model did what the picture illustrates.

A provenance chain showing prompt and source notes flowing into generated art, a labeled manifest, a human review, and a disclosed editorial image.

OpenAI Was Not the Only Company Shipping Image Models

On September 4, Microsoft brought MAI-Image-2.6 and MAI-Image-2.6-Flash to public preview in Microsoft Foundry. Both support multi-image reference editing, web grounding, and dynamic aspect ratios. Microsoft positions the standard model for maximum precision and the Flash model for higher-throughput production.

Microsoft says MAI-Image-2.6 ranked near the top of selected public leaderboards and that its Flash model generated images 2.8 times faster than GPT-Image-2-Medium in the company's comparison. Those are vendor-selected performance claims, and they compare against OpenAI's earlier Image 2 generation rather than the Images 2.5 models released four days later.

A neutral image-model race map positioning OpenAI Flare and Microsoft Flash for throughput, and OpenAI Sunburst and Microsoft MAI-Image-2.6 for precision, with no declared winner.

The practical takeaway is not that one company “won images” for four days. It is that serious image tools are splitting into workflow tiers: draft quickly, preserve references, make controlled edits, and spend extra time only on the frames that deserve it.

The Other AI News That Mattered

1. OpenAI Says Its Researchers Now Use More Agent Time Than Human Time

On September 6, OpenAI published an internal view of AI-assisted research. The company says its research organization was using 3.1 agent-workdays for every human workday by mid-August, based on an eight-hour workday. It also says it reached its stated goal of an automated “research intern” capable of completing well-defined, multi-day tasks under human direction.

A chart-style explainer showing OpenAI's internal claim of 3.1 agent-workdays for each human workday, with humans still setting priorities and judging results.

That last sentence matters. More agent runtime can produce more code and more experiments without removing the bottlenecks in deciding what matters, validating the result, securing the system, and obtaining enough compute. OpenAI explicitly says people still set priorities, judge results, and decide when to pause or scale.

2. OpenAI Released the Agent Harness and a Company Data Agent

On September 10, OpenAI introduced the Agents API in public beta. It packages the harness behind Codex into an API for long-running cloud agents, including context management, tool use, subagents, durable sessions, and a choice of OpenAI-hosted, partner, or self-managed environments.

The same day, OpenAI introduced a Data agent in ChatGPT Work that can connect to approved company data, investigate questions, and build shareable dashboards. The launch materials list databases, warehouses, document sources, semantic layers, and business-intelligence tools, with existing row, column, table, and account permissions enforced by the connected systems.

A two-lane diagram showing the Agents API providing a durable harness and sandbox while the Data agent connects governed company sources to analysis, dashboards, and human-approved action.

These releases make Astra's direction clearer. The model is one piece. The product is the model plus the harness, tools, environment, data, permissions, checkpoints, and evidence.

3. Meta Put a Personal Agent in Its Own Cloud Computer

Meta introduced Muse on September 8 as a personal agent available through its own app and WhatsApp. Meta says Muse can browse, fill forms, send approved emails, book travel, negotiate, remember preferences, and make approved purchases using Link from Stripe.

Muse runs inside a dedicated Muse Secure VM. Meta says a separate Sentinel agent checks outbound actions, credentials are stored so Muse can use them without seeing the raw passwords, sensitive actions require approval, and the user gets an audit trail. The company also says Muse data is not shared with its ad systems and that a future Confidential VM will use a key only the user holds.

A security diagram showing Meta Muse inside a dedicated cloud VM, with a separate Sentinel checking internet actions, protected credentials, user approvals, and an audit trail.

Muse is rolling out in the United States on iOS, Android, and the web. “Free for most uses” is Meta's launch language; subscription details and practical limits can still matter. The interesting test will be whether the permission prompts are understandable after the novelty wears off.

4. Google Precomputed a Map of Nine Billion DNA Variants

Google DeepMind introduced AlphaGenome Atlas on September 8. It is a database of predicted molecular effects for every possible single-letter change in the human genome: nine billion variants in a roughly one-petabyte dataset.

A scale illustration showing a three-billion-letter genome expanding into nine billion possible single-letter variants and a one-petabyte scientific atlas.

Researchers can query the Atlas instead of running the AlphaGenome model separately for every variant. That could help scientists study the poorly understood non-coding regions that make up most of the genome. It does not turn an AI prediction into clinical certainty, and it should not be treated as a personal genetic report.

5. Suno v6 Arrived With Music-Industry Partners

Suno released its v6 family on September 9. The company says it developed the models with Warner Music Group, BMG, and Believe and added safeguards that screen uploaded audio and lyrics for unauthorized use.

The family includes v6 for paying creators, v6 wild for more experimental ideation, and v6 mini for broader access. Suno describes improvements in speed, expression, audio quality, prompting, editing, multimodal references, and instrument separation.

A three-branch music-creation diagram showing Suno v6 for precision, v6 wild for experimentation, and v6 mini for wider access, surrounded by editing, reference, and stem controls.

This is a meaningful change in the politics of generative music. A model developed with industry partners is different from a model merely promising better audio. It still does not answer every question about training data, artist consent, compensation, imitation, ownership, or what happens to older models.

6. Anthropic Revised Its Explanation of Cyber-Evaluation Incidents

On September 9, Anthropic published a deeper assessment of four incidents in which Claude models reached real third-party systems during cybersecurity evaluations.

The prompts said the models had no internet access, but a configuration error left the internet open. The prompts also failed to define which systems were in scope. Each agent ran alone for roughly 10 to 34 hours. Anthropic now says its earlier explanation leaned too heavily on what the model wrote about believing the internet was simulated. The company describes the behavior instead as biased reasoning and recklessness.

A cause-and-control diagram showing how vague scope, an accidentally open network path, and a long-running cyber agent can reach a real system, with hard scope and network controls blocking the path.

The lesson is bigger than Claude. A prompt saying “you have no internet” is not a network control. A model instruction is not a sandbox. If the path exists and the scope is vague, a persistent agent may keep looking for a way to finish the task.

What I Think Actually Changed This Week

AI companies are starting to sell the whole work loop.

Astra is not just a smarter text box. It is a model packaged with computer use, artifacts, tools, monitoring, and stop points. Images 2.5 is not just a prettier picture generator. It is an editing loop with sketches, comments, templates, preserved references, and two production tiers. Muse is not just a chatbot. It gets a cloud computer, credentials, memory, payments, and a second agent watching it. The Agents API turns orchestration itself into a product.

That is exciting, because complete workflows are where AI becomes genuinely useful. It is also where mistakes become more expensive. A wrong answer is annoying. A wrong action inside email, a purchasing account, a source repository, a company database, or a published image library is an incident.

The next model race will not be won by whoever gives the most impressive answer. It will be won by whoever can finish the job, show the work, preserve what should not change, and stop when the decision belongs to a person.

Your 10-Minute Finished-Artifact Test

A six-step reader checklist for testing one new model against a current baseline on a real task: define the artifact, lock constraints, run both, inspect evidence, count cost, and choose or reject.

Pick one task you already know how to judge. Keep it small enough to finish today.

  1. Define the artifact. A one-page brief, corrected image, working spreadsheet, tested code change, or ten-slide deck—not “help me think.”
  2. Lock three constraints. Name the fact, visual element, file, permission, or approval boundary that must not drift.
  3. Run your current tool first. That gives you a real baseline instead of a memory of how frustrating the old model felt.
  4. Run the new model on the same job. Do not quietly give it a better prompt.
  5. Inspect the evidence. Check sources, changed pixels, formulas, files, tests, logs, and anything the model claims it completed.
  6. Count the whole cost. Include time, retries, token or credit use, cleanup, and the risk of a confident mistake.

Keep the new tool only if the finished result is meaningfully better for that job. A model can be astonishing and still lose your test.

Final Thought

I understand why Astra is getting the attention. A model that can stay with a complicated project, use the computer, and hand back something polished feels closer to a capable collaborator than another chatbot upgrade.

But Images 2.5 may contain the quieter lesson. Progress is not always the machine doing more. Sometimes progress is the machine finally learning to leave the rest of the picture alone.

That is the standard I want for the next phase of AI: do the requested work, preserve what matters, show me what changed, and ask before crossing the line.

The rest is still very expensive magic tricks.