Developers and organizations using Together GPU Clusters can now run workloads at significantly lower cost using preemptible compute with a 5-minute drain window, expanding options for cost-sensitive use cases.
Evidence
+Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.
Llama Guard 4 12B model removed from Together AI pricing page
Users relying on Llama Guard 4 12B ($0.20/token) through Together AI will need to find alternative providers or models for content moderation/safety use cases.
Together AI published benchmark comparison of GLM-5.3 vs GLM-5.3 Flash models on DeepSWE
Developers can now evaluate the cost-performance tradeoff between GLM-5.3 and its faster Flash variant, with Flash offering 17x lower cost at the expense of ~5.6 points in pass@1 performance.
Evidence
+GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing
We ran 900 DeepSWE rollouts on GLM-5.3 and GLM-5.3 Flash. Flash gives up 5.6 points of pass@1 at 17x lower cost, and only 2.6 points at pass@4.
New DeepSeek V4 Pro 0813 model introduced with different pricing
Users now have access to a new DeepSeek V4 Pro 0813 variant with lower input pricing ($1.32 vs $1.74) and higher output pricing ($3.96 vs $3.48), with reduced cached token pricing ($0.13 vs $0.20).
New benchmark comparing DeepSeek-V4 Flash vs GPT-5.6 Luna on coding performance and cost efficiency
Developers can now compare cost-effectiveness of different LLMs for coding tasks; highlights DeepSeek's superior price performance for coding workloads
Evidence
+DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding
We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
Added DeepSeek V4 Flash 0731 model with pricing and cache support
Users can now access DeepSeek V4 Flash 0731 model through Together AI's platform with input cost of $0.14/MTok, cached input cost of $0.03/MTok, and output cost of $0.28/MTok
New blog post about Kimi K3 vs Claude Fable 5 performance and cost comparison
Users can now access comparative analysis of Kimi K3 and Claude Fable 5 models on DeepSWE benchmark, helping inform model selection decisions based on pass rates and cost-efficiency metrics.
Evidence
+Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding
We ran 452 DeepSWE rollouts on Kimi K3 and Claude Fable 5. Fable leads pass@1 by 1.4 points; Kimi K3 wins pass@4 and delivers 2.8x the solves per dollar.