Skip to content

Daily briefing

Model news

Daily headlines around the freshest and most-listed models in the catalog, fetched via Tavily. Also available as RSS or JSON.

Abacus.AI Smaug Models Cut Agent Costs 100x [2026]

Smaug Mini is the smallest of the three, based on Qwen3.8 27B and aimed at compact multimodal tasks where a lighter model is cheaper to run at scale. Abacus.AI describes Smaug Mini as suited to smaller, high-volume jobs rather than the deep multi-step reasoning chains that Smaug Agentic is built for. Between the three, Abacus.AI is effectively offering a size ladder: pick Mini for cheap, high-volume multimodal tasks, Flash for always-on agents with long context, and Agentic for the hardest [...] According to Abacus.AI’s own materials on its open-source page, the underlying methodology combines human-curated, real-world agentic traces with synthetic data grounded in hard examples, then applies that single training recipe across three different open-weight bases: Smaug Flash on DeepSeek V4 Flash, Smaug Mini on Qwen3.8 27B, and Smaug Agentic on Kimi K3. That’s a notable design choice: rather than building one model and shrinking it, Abacus.AI runs the same fine-tuning process across three [...] | Model | Base model | Parameter scale | Primary use case | --- --- | | Smaug Agentic | Kimi K3 | 2 trillion | Complex coding, long-running agentic loops | | Smaug Flash | DeepSeek V4 Flash | Not disclosed | Always-on enterprise agents, long context, heavy tool use | | Smaug Mini | Qwen3.8 27B | 27 billion | Compact multimodal tasks, high-volume jobs | ## The Fine-Tuning Technique Behind the 15-20% Gain

shattered.io· 1d ago·Qwen3.8 27B
The TechBeat: AI Coding Tip 035 - Split Every Skill Description Into Three Sentences (9/12/2026)

By @noufalb [ 13 Min read ] AI is removing routine junior work, but those tasks also helped build expertise. Companies may be trading short-term productivity for long-term capability debt. Read More. ## Qwen3.8-27B Cold Fusion Cuts Thinking Tokens Without Sacrificing Performance TechBeat's image-308238 By @aimodels44 [ 9 Min read ] Explore Qwen3.8-27B Cold Fusion, a 27B AI model designed to cut thinking tokens while retaining strong quantized reasoning performance. Read More.

hackernoon.com· 2d ago·Qwen3.8 27B
DeepSeek Releases V4.1 Flash, a 748-Billion Flash That Now Sees

#### Qwen3.8-Flash-Next goes open source and revives DeepSeek's n-gram idea DeepSeek Launches V4-Flash-Vision-Exp: Cheap Model Now Sees Images, Takes On Opus 4.8 #### DeepSeek Launches V4-Flash-Vision-Exp: Cheap Model Now Sees Images, Takes On Opus 4.8 MiniMax M3: the Chinese Open-Weights Model Taking On GPT-5.5 #### MiniMax M3: the Chinese Open-Weights Model Taking On GPT-5.5 IFM Releases K2 Horizon, Six Open Models From 0.9B to 375B With the Training Data [...] The number circulating most, the 552 billion, is only the skeleton. On top goes an n-gram of roughly 200 billion, bringing the real count to 748 billion, the same technique we saw relaunched with the n-grams of Qwen3.8-Flash-Next. Forum users mocking the "half-terabyte Flash" caught the right paradox. The model is built to run fast and cost little in production, not to sit comfortably on a laptop. [...] ## Conclusions V4.1 Flash is a frontier multimodal model with MIT weights, matching Opus 5 on agentic coding and costing a fraction of the price. The price to pay is bulk. Whoever has the 512 GB brings it home and works offline, everyone else goes through the discounted API. The Hangzhou house keeps making the same move, raising the bar and lowering the bill. Enjoyed the article? ### Related Articles Qwen3.8-Flash-Next goes open source and revives DeepSeek's n-gram idea

pasqualepillitteri.it· 2d ago·Qwen3.8 Flash Next
I Built a 256GB “E-Waste” AI Rig to Escape Token Fees. Here’s What Nobody Tells You.

Configuration: dual X99, two E5-2696 v4, 128GB DDR4, four CMP 170HX unlocked to 64GB each, 256GB VRAM total. Qwen3.8 Flash-Next official native version, TP2\PP2, 1M context. About 60 tok/s single stream, about 450 tok/s at 16 concurrent. About 11 kWh per day. It is now my main machine. It writes code and runs tasks with almost no cache hit rate. I have deployed DSV4F, DSV4F Exp, and Qwen3.8 27B locally. I kept Qwen3.8 Flash. On my tasks, it beats DSV4F and GLM5.3 Flash. [...] Know the boundary. Qwen3.8 Flash is a fast typist, not a universal brain. Complex logic, long-chain reasoning, and cross-domain knowledge integration go to cloud API. Local handles frequent, private, cost-sensitive work. Cloud handles rare, hard, high-value work. Squeeze VRAM. With 256GB VRAM, deploy speculative decoding or Medusa heads. Use idle VRAM for a 2B draft model. Single-stream speed jumps to 90+ tok/s. Free performance. [...] Qwen3.8 Flash balances capability, speed, VRAM, power, and engineering complexity. Token freedom has three layers. Physical freedom: you have VRAM, compute, and concurrency. The model runs and serves multiple requests. 256GB VRAM, TP2\PP2. 60 single-stream, 450 at 16 concurrent. That is productivity.

vocal.media· 4d ago·Qwen3.8 Flash Next
LattePanda unveils micro compute module delivering 115 TOPS for AI

In testing, LattePanda Mu Ultra achieved text generation speeds of 18 tokens/s with Qwen3.5-9B and 55 tokens/s with Qwen3.5-2B, using INT4 quantization with OpenVINO GenAI on the iGPU. These results demonstrate its ability to support responsive, local conversational AI without cloud inference. This makes it suitable for applications such as local voice assistants, offline document processing, and private knowledge retrieval, keeping sensitive data on the device. [...] Display Output: - Up to 3 × HDMI / DisplayPort - 1 × eDP Dimensions: - 69.6 × 60mm [...] Upgrade with a Modular Design LattePanda Mu Ultra retains the standardized form factor and connector of the LattePanda Mu family and is largely compatible with existing carrier boards. When paired with the 3.5-inch Lite Carrier Board, Mini Carrier, or M.2 M Key Carrier Board, it enables users to upgrade existing systems without redesigning the entire carrier board.

vir.com.vn· 4d ago·Qwen3.5 9B
GPT-6 Astra: A new generation of intelligence

GPT‑6 Astra is our best model for adhering to existing templates and producing slides that are well laid out and succinctly convey key points with a structured narrative. It creates clear, well-structured documents, presentations, spreadsheets, and analyses that follow your templates and match your writing and visual style. Astra is also trained to specifically pull only the context that matters into outputs, instead of repeating information unnecessary for the work at hand. All this means it [...] Astra is our most aligned model, with substantial improvements in understanding user intent and model behavior—you can delegate tasks with greater confidence in Astra’s judgment. As one way that we test this, we built a new evaluation informed by the Hugging Face incident that evaluates whether a model facing a difficult or impossible task will go beyond its intended scope. Compared to GPT‑5.6 Sol, which without production safeguards went beyond the authorized target 48% of the time, GPT‑6 [...] ## Aligning and deploying GPT‑6 Astra responsibly Astra is our most aligned model. Astra excels at exercising care, respecting task boundaries, and communicating transparently. This work is the latest product of our long-running research program focused on training models that remain aligned with human intent from start to finish.

openai.com· 4d ago·GPT-6 Astra+3 outlets
AI & Robotics in 2026: 16 Recent Developments for Investors

## 5. Alibaba expanded developer options with open Qwen weights THNQ index constituent Alibaba Group Holding (BABA), the commerce and cloud-computing company, released Qwen3.8-Flash-Next in August 2026. It made the model’s weights, the numerical settings learned during training, available for developers to deploy and adapt under its license. [...] Faster inference can shorten the repeated cycles agents use to code, research, call tools, and complete complex tasks. Taken alongside Qwen’s more efficient architecture, GLM’s lower-cost open models, and OpenAI’s custom Jalapeño processor, the Cerebras partnership shows how broadly the industry is attacking inference economics. Improvements across model design, silicon, and computing architecture could expand the range of AI workloads that become economically practical. [...] AI agents are becoming more persistent and connected. Hark partnered with NVIDIA, while private payments company Stripe agreed to acquire OpenRouter, connecting payments with access to more than 400 AI models. Inference is becoming a systems-level competition. Qwen, GLM, OpenAI/Broadcom, and Cerebras are attacking cost and latency through model architecture, open weights, custom silicon, and alternative computing platforms.

etfdb.com· 4d ago·Qwen3.8 Flash Next
Run Qwen 3.8 Locally: Ultimate RTX 5090 Guide (2026)

Qwen 3.8-27B ships under an Apache 2.0 license, Alibaba’s usual choice for its open releases, which means it’s free to use commercially with no restrictive clauses to navigate. Combined with native image and video understanding baked into the same 27B checkpoint, that makes it a genuinely capable local alternative to closed, subscription-gated multimodal APIs. You can review the full model card, weights, and benchmark details on Qwen’s official Hugging Face page. [...] Qwen 3.8 is the newest generation in Alibaba’s open-model family, and it ships in two very different tiers. There is a massive 2.4T-parameter “Max”-class checkpoint that is cloud-only and never intended for consumer hardware, and there is Qwen 3.8-27B, a dense, deployment-friendly model that is fully open-weight under an Apache 2.0 license and built specifically to run on local hardware like a single RTX 5090. When people talk about running Qwen 3.8 locally, the 27B release is the one that [...] If you’d rather have a graphical chat interface instead of the command line, LM Studio is the most polished option and it lists Qwen 3.8 directly in its model catalogue. Open LM Studio, search for “Qwen3.8” in the Discover tab, choose the 27B GGUF build sized for your VRAM, download it, and load it from the My Models tab. LM Studio also exposes a local OpenAI-compatible server, so you can point existing apps at it exactly the way you would an OpenAI endpoint.

technosports.co.in· 5d ago·Qwen3.8 27B+1 outlet
China’s AI models aren’t proof that US chip controls have failed

## Share ## Share Despite stringent US export controls on chips and semiconductor exports to China since 2022, the latest Chinese artificial intelligence (AI) models — Kimi K3, DeepSeek V4 Pro and GLM-5.3 — are approaching or matching leading US models on many public benchmarks, often at lower costs and with ‘open weights’ that allow more user modification. Some take this as evidence that US export controls must have failed.

eastasiaforum.org· 5d ago·DeepSeek V4 Pro
Can You Trust These 7 AI Models? Their Worst Weeks Answer It

Moonshot released Kimi K2 Thinking in November 2025 and Kimi K3 on July 16, 2026. ### Current Status Kimi is the most popular Chinese AI tool inside Western enterprises, often used without IT approval because it is free and runs in a browser. Moonshot open-sourced K3's weights on July 27 and is negotiating hosting deals with Microsoft, Amazon, and Google. ### Top Controversies [...] The model still censors politically sensitive topics, and self-hosting leaves that intact. A company that ships open weights and closes ranks on basic data questions is trading away brand trust it could keep. ## Moonshot AI's Kimi ### Brief Backstory Kimi comes from Chinese AI lab Moonshot AI, backed by more than $1 billion from Alibaba and Tencent. DeepSeek got the bigger headlines in early 2025, and Kimi has since become the dominant Chinese AI tool by usage. [...] Incogni's 2026 Gen AI and LLM privacy ranking found that Moonshot lets users opt out of training data use without making the process clear. Users have to email support and verify their identity manually, with no in-product toggle. Kimi's open weights make a U.S. ban hard to enforce, since the model runs independently of Moonshot's servers. ### Public Reaction and Company Response Moonshot's privacy policy acknowledges an opt-out right and puts the entire burden of using it on the user.

news.designrush.com· 5d ago·Kimi K2 Thinking
Take on your most ambitious work with GPT-6 Astra ...

Today, GPT-6 Astra, the latest and most capable OpenAI model, is generally available on Amazon Bedrock. You can call the model directly through the Amazon Bedrock APIs or configure ChatGPT Work and Codex to use GPT-6 Astra on Amazon Bedrock. As part of this launch, OpenAI is also introducing new enterprise plugins for ChatGPT Work that extend Astra’s browser-use capabilities across common business applications. The Amazon Bedrock inference engine delivers the scalability and reliability [...] Model-level safeguards work alongside the security and governance controls of Amazon Bedrock. OpenAI evaluated GPT-6 Astra through its Preparedness Framework, which assesses model capabilities across safety-relevant domains and applies progressively stronger safeguards as those capabilities advance. It’s the first OpenAI model to reach the Critical classification for cybersecurity capability. At this level, automated safeguards monitor misuse in real time and can pause or stop activity that [...] Organizations are already running AI agents that write code, analyze data, and automate complex workflows at production scale on Amazon Bedrock. GPT-6 Astra raises the potential of what those agents can deliver. It applies deeper reasoning and sharper judgment to complex business decisions, works across software and files, and produces professional-quality output aligned with organizational voice, templates, and standards.

aws.amazon.com· 5d ago·GPT-6 Astra
America Must Protect Its Training Data

And China’s labs have taken advantage. Recent reporting puts the top six Chinese labs’ annual spending with U.S. AI data companies at over $500 million. And experts have long speculated as much: SemiAnalysis wrote in January that Surge sells training environments “to Chinese labs like Moonshot and Z.ai,” and that access “played a huge part in increasing the capabilities for Kimi K2 Thinking and GLM-4.6.” [...] Kimi, like all AI models, was trained using two key resources: semiconductor chips and training data. Policymakers have focused heavily on limiting China’s access to the U.S.’s best chips— producers such as Nvidia and Intel face myriad export controls restricting sales to China. [...] # America Must Protect Its Training Data Washington has spent the summer alarmed by Chinese artificial intelligence (AI). After the release of Kimi K3—a Chinese model that pulled within reach of the U.S. frontier models—the White House began weighing plans to ban enterprise use of Chinese models, and Congress has escalated probes into the use of Chinese AI by U.S. companies.

lawfaremedia.org· 6d ago·Kimi K2 Thinking
Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses

∙ Paid Avid Artifacts readers know that we have been covering not only models but also their licenses for quite some time. There was a period when custom licenses were all the rage, for example the custom Qwen2.5 72B-Instruct license or the Llama licenses. DeepSeek had a custom license for DeepSeek V3 before R1 changed it to MIT, which has resulted in many (Chinese) model makers adopting MIT or Apache 2.0 licenses in 2025. [...] View more details on all the models in this issue at our Artifacts Hub. Visit artifactshub.ai ### Models #### General Purpose [...] > If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 10 billion US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must pass Z.AI’s security review before using the Software or its derivative works for any commercial purpose. The scope and method of the security review shall be reasonably determined by Z.AI.

interconnects.ai· 6d ago·Qwen3.8 Flash Next
Gemini 3.8 Flash vs GLM 5.3 Flash: Benchmarks, Pricing & Speed (2026)

GLM 5.3 Flash itself is a paid model, but Z.AI offers GLM-4.7-Flash and GLM-4.5-Flash completely free for all registered users. For GLM 5.3 Flash specifically, the GLM Coding Plan starts at approximately $18/month, and off-peak hours (including all weekends) bill at half the standard credit rate. ### Which model is faster? [...] A notable detail: GLM-4.7-Flash and GLM-4.5-Flash are completely free for all registered Z.AI users, with no per-token charge. GLM 5.3 Flash itself is paid, but the Coding Plan gives it 3x the usage quota compared to the GLM-5.3 flagship. Off-peak hours (including all weekend hours) consume only 50% of standard credits on Coding Plans. Additionally, the GLM-5.3-Flash Usage Campaign (running until September 20) gives paid plan users unlimited GLM-5.3-Flash usage via ZCode daily from 23:00 to [...] | Weights | Closed | Open (MIT license) | | Free Tier | Google AI Studio free tier | GLM-4.7-Flash is fully free |

intelligentliving.co· 6d ago·GLM-4.5
Reports indicate that DeepSeek plans to build a data center in Inner Mongolia and install 160,000 Huawei AI chips, accelerating the development and operation of AI using Chinese-made GPUs.

DeepSeek is a company that develops AI models with performance comparable to cutting-edge AI from OpenAI and Anthropic. DeepSeek is characterized by its open-source approach, releasing its AI models free of charge. On August 31, 2026, it released DeepSeek-V4-Flash-Vision-Exp, which boasts image recognition performance equivalent to Claude Opus 4.8. [...] DeepSeek has released 'DeepSeek-V4-Flash-Vision-Exp,' an open model with image recognition capabilities, achieving performance equivalent to Claude Opus 4.8 in tasks including image recognition - GIGAZINE [...] Huawei releases 'Pangu Pro MoE 72B', a language model trained in China's AI ecosystem, and open-sources inference technology NVIDIA announces its most powerful open-source model in the US, the 'Nemotron 3 Ultra,' and reports the start of mass production of its AI server, 'Vera Rubin.' Qwen3.8, a Chinese AI with 2.4 trillion parameters, has been released as an open model for free, boasting performance comparable to the most advanced models from OpenAI and Anthropic.

gigazine.net· 6d ago·DeepSeek V4 Flash Vision Exp
Alibaba releases Qwen3.8 Flash Next, the local Qwen4 preview

& Tutorials 2 Guides & Tutorials AI for Professionals 1 AI for Professionals Reports & Analysis 2 Reports & Analysis Finance 1 Finance Agent Infrastructure 1 Agent Infrastructure [...] · 7 min read Claude Code self-hosted keeps your code in-house, but not the AI model Claude Code & Anthropic #### Claude Code self-hosted keeps your code in-house, but not the AI model · 10 min read Astra is late to paid plans, OpenAI hands out a reset for every day of the wait AI News & Trends #### Astra is late to paid plans, OpenAI hands out a reset for every day of the wait · 7 min read IFM Releases K2 Horizon, Six Open Models From 0.9B to 375B With the Training Data [...] Most Read AI News & Trends 20 AI News & Trends Videogiochi 1 Videogiochi Cybersecurity 20 Cybersecurity Google AI & Gemini 12 Google AI & Gemini Apple 4 Apple Automotive Tech 2 Automotive Tech Diritto & Normative 2 Diritto & Normative Claude Code & Anthropic 4 Claude Code & Anthropic Prompt Engineering 1 Prompt Engineering Benchmarks & Comparisons 1 Benchmarks & Comparisons Intelligenza Artificiale 1 Intelligenza Artificiale Guides & Tutorials 2

pasqualepillitteri.it· 7d ago·Qwen3.8 Flash Next