Date ▾ Name Lab Type Stars Downloads Citations Params Active Tokens Context Intel Open Questions Tasks
2026-09-24 KREX: Concurrent Kernel Benchmarking on Shared GPUs via Region-Granular Exclusivity HKUST, Alibaba paper
2026-09-22 Rigel Base UC Berkeley model 2.3B360M294.91K
2026-09-22 Agensh: Scaling Organizational Intelligence to 1,024 Agents Microsoft paper
2026-09-21 1% of Tokens Can Be Enough: Gradient Estimation in On-Policy Distillation MBZUAI paper
2026-09-21 ★ Grok 4.7 SpaceXAI model 500K46
2026-09-21 MiMo-V2.6-Distill-Qwen-9B Xiaomi model
2026-09-21 MiMo-V2.6-Flash Xiaomi model 309B15B48T1M
2026-09-21 ★ MiMo-V2.6-Pro Xiaomi model 1.02T42B30T1M46
2026-09-21 Alibaba Appoints Dayiheng Liu Head of Qwen LLM Team After Months of Reorganization; Jingren Zhou Moved to Chief Scientist in June The Information Alibaba news
2026-09-21 SpaceXAI Releases Grok 4.7: New Larger Base Model, Longer-Horizon RL, Same $2/$6 Pricing, AAII v4.3 46 SiliconANGLE SpaceXAI news
2026-09-21 Xiaomi Releases MiMo-V2.6: Pro (1T/42B, MIT) Debuts as Top Open-Weights Model on AAII v4.3 (46), Plus Flash, Distill-Qwen-9B, and Open RL Environments Xiaomi MiMo Xiaomi news
2026-09-20 ★ Step 5 Preview StepFun model 600B27B1M44
2026-09-20 StepFun Launches Step 5 Preview: 600B/27B MoE, 1M Context, $1/$2.70 API, AAII v4.3 44; Open Weights Promised for October 15 StepFun StepFun news
2026-09-19 DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Agentic Training at Scale DeepSeek paper
2026-09-18 ★ Nemotron-SEA-LION-v4.8 AI Singapore model 120B12B262.14K
2026-09-18 Uplifting AI in Southeast Asia: Announcing Nemotron-SEA-LION-v4.8 in Collaboration With NVIDIA (30B-A3B and 120B-A12B, MIT) AI Singapore AI Singapore news
2026-09-17 Motif 3 API Free Until September 29 as Motif Technologies Courts Developers After the Sovereign-Model Elimination AI Times Motif Technologies news
2026-09-17 Nex-N2.5-Pro Weights Published (397B, Apache 2.0, BF16 and FP8) Hugging Face Nex-AGI news
2026-09-17 Xiaomi Livestreams MiMo-V2.6 Flash and Pro RL Training With a Public Dashboard; Trial Models Named MiMo-X-Pro-Preview and MiMo-X-Flash-Preview Xiaomi MiMo Xiaomi news
2026-09-17 Zhipu Reports a GLM-5.3-Driven Infrastructure Agent That Lifted End-to-End Throughput 3.2x on a 100K+ Domestic-Chip Cluster in Under Two Weeks QbitAI Z.ai news
2026-09-16 Agora NVIDIA paper
2026-09-16 Arcee AI Raises a Series B Led by Vista, Cambium, and Emergence at a Valuation Above $1B; Says Its 2025 Lineup Including Trinity Large Was Built for ~$20M Arcee AI Arcee news
2026-09-16 Cohere and Aleph Alpha Sign Definitive Business-Combination Agreement: Dual HQ Berlin and Toronto, 1,000+ Staff, Close Expected Later in 2026 Cohere Cohere news
2026-09-16 Google Launches the DeepMind Institute, a Founder-Led AGI Policy Vehicle Directed by Legg, Manyika, and Hassabis DeepMind Institute Google news
2026-09-16 Agora: Git as Shared Memory for Collective AutoResearch, 13 Agents Over 12 Days Producing 165 Verified Claims arXiv NVIDIA news
2026-09-16 OpenAI Publishes a Model Misalignment Reporting Framework and Six Incident Reports, Including an Unreleased Astra-Family Model Writing Jailbreak Instructions Into Its Own Compaction Summaries OpenAI OpenAI news
2026-09-16 OpenAI in Talks for New Funding at a $1.2T Valuation Sought by Investors and $1.5T+ Sought by the Company (Not Closed) The New York Times OpenAI news
2026-09-16 PFN and Mitsubishi Heavy Industries Form a ¥10B Capital Alliance Preferred Networks PFN news
2026-09-15 Gemini 3.8 Live Google model
2026-09-15 Open-SWE-Traces NVIDIA dataset
2026-09-15 Salesforce Koa Salesforce model 120B
2026-09-15 LimiX-2 Tsinghua model 400M
2026-09-15 Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking: Native Audio Models in 97 Languages; Extended Thinking Takes #1 on the AA Speech-to-Speech Index (82.6) Google Google news
2026-09-15 Standalone Kanana App to Shut Down October 15 as Kanana Moves Into KakaoTalk and KakaoMap; Roadmap Due at if(kakao)26 AI Times Kakao news
2026-09-15 Elsevier Deploys LG AI Research's Chemistry Vision Model in Reaxys to Extract Structures and Reactions From Scientific Images Elsevier LG news
2026-09-15 Open-SWE-Traces v1.2: 511,668 Agentic SWE Trajectories (CC-BY-4.0) With a Benchmark-Exploit Audit Paper Hugging Face NVIDIA news
2026-09-15 Salesforce Koa: First Salesforce-Branded LLM, Post-Trained From Nemotron-3-Super-120B With GRPO for Agentforce (Closed Weights, GA Winter 2026) Salesforce Salesforce news
2026-09-14 Stellar Colosseum Google paper
2026-09-14 METIS Peking University model
2026-09-14 Siri AI Launches on Next-Generation Apple Foundation Models 'Custom-Built in Collaboration With Google and Its Gemini Models' (English Beta) Apple Apple news
2026-09-14 LG AI Research Showcases EXAONE Discovery, EXAONE Omni-Inspect, and EXAONE Business Intelligence as Industrial 'Expert AI' Products The Korea Herald LG news
2026-09-14 Microsoft AI Publishes a Draft MAI Code of Conduct: Models Must Not Resist Shutdown, Widen Scope, or Hide Reasoning (Six-Week Consultation) Microsoft AI Microsoft news
2026-09-13 Flattening Every Memory Peak in Long-Context MoE Training Salesforce paper
2026-09-13 Dream-RSI University of Maryland paper
2026-09-13 Intern-S2-397B Release Weights Published Under Apache 2.0, Succeeding the July Preview Hugging Face PJLab news
2026-09-13 Upstage Opens a Free Preview of Solar Mini 4, a 35B-A3B Small Model Said to Match Solar Pro 4 at One-Seventh the Size; API GA Planned for September 22 AI Times Upstage news
2026-09-12 Tabby Huawei model 145M
2026-09-12 StepAudio 3 Realtime StepFun model
2026-09-12 Tabby: Noah's Ark Lab Releases a 145M Time-Series Foundation Model With Its Complete Open Pretraining Recipe arXiv Huawei news
2026-09-12 StepAudio 3 Realtime and StepAudio 3 Gen Technical Reports: Think-While-Speaking Duplex Audio LM and a Unified Discrete Audio Generator arXiv StepFun news
2026-09-11 Never Give Up Mila, Ai2 paper
2026-09-11 Agnes 3.0 Flash Sapiens AI model 33B1M36
2026-09-11 Atria Dawn Preview PJLab model 744B40B262.14K
2026-09-11 StepAudio 3 Music StepFun model
2026-09-11 SAS (Simple Attention Sparsification) Tencent paper
2026-09-11 ★ ZGCM-1 ZGCA model 7B4.19T262.14K
2026-09-11 Cohere in Advanced Talks to Raise $2–3B at a $20B Valuation (Globe and Mail); Series E Not Yet Finalized PYMNTS Cohere news
2026-09-11 In Depth: Inside the $50 Billion Rise of Moonshot AI — Pre-Money Valuation Quintupled in Six Months After Kimi K3; G-Round Still Open Caixin Global Moonshot AI news
2026-09-11 Kimi K2.8 Preview Rolls Out to All Kimi Code Members With 1M Context and Low/High/Max Effort Levels; No Weights or Report Kimi Code docs Moonshot AI news
2026-09-11 Agnes 3.0 Flash Listed on Artificial Analysis at 36 (v4.3); Open-Weight 33B Preview Checkpoint Released Under Apache 2.0 Hugging Face Sapiens AI news
2026-09-11 Atria Dawn Preview Released: Agentic Model Post-Trained From GLM-5.2 (744B MoE, MIT, 256K) With a 143-Author Paper on a Verifiable Experience Pipeline Hugging Face PJLab news
2026-09-11 StepAudio 3 Music Technical Report: MoE Planner Plus Flow-Matching DiT for Long-Form Music, Elo 1105 on the AA Music Arena Vocals Leaderboard arXiv StepFun news
2026-09-11 Falcon-OCR v1.5: GRPO-Post-Trained 270M Early-Fusion VLM Reaches 93.15 on OmniDocBench (Pipeline) Hugging Face TII news
2026-09-11 ZGCM-1: Zhongguancun Academy Releases a Fully Open 7B Model Pretrained From Scratch on 4.19T Tokens (MIT), With Pre-Training and Mid-Training Checkpoints arXiv ZGCA news
2026-09-10 ACE2S-SHiELD+ Ai2 model
2026-09-10 Occamy-1.0 Alibaba model 35B3B
2026-09-10 OmniTable Ant Group paper
2026-09-10 North Small Translate 1.0 Cohere model 218B25B16.38K
2026-09-10 deepseek-recipe DeepSeek library
2026-09-10 ★ DeepSeek-V4.1-Flash DeepSeek model 552B16B45T1M3944/100
2026-09-10 NASA-IBM Lunar Foundation Model IBM model
2026-09-10 Magenta MBZUAI paper
2026-09-10 ★ Agents API OpenAI announcement
2026-09-10 SenseNova-U1.5 SenseTime model 8B
2026-09-10 Intern Lumina U2 PJLab model 16B1B
2026-09-10 Mixtures-of-Experts Overfit More to Repeated Data Stanford paper
2026-09-10 T1 Tencent model 122B10B
2026-09-10 ACE2S-SHiELD+ Climate Emulator Weights Released (Apache 2.0), Trained to Separate SST and CO2 Forcing Hugging Face Ai2 news
2026-09-10 Occamy-1.0: Alibaba's Accio Team Releases a Cost-Efficient 35B-A3B Co-Work Model Further Trained From Qwen3.6-35B-A3B arXiv Alibaba news
2026-09-10 Anthropic Publishes an Evaluation of Tactical Intelligence Targeting and Conventional-Weapons Capabilities Across Frontier Models Anthropic Anthropic news
2026-09-10 Threat Intelligence Report, September 2026: Seven Threat Groups, 50+ Organisations, an 8,913-Article Influence Operation Anthropic Anthropic news
2026-09-10 North Small Translate 1.0: Open-Weight 218B-A25B Translation Model on the Command A Plus Base, WMT26 83.60 (CC-BY-NC-4.0) Cohere Cohere news
2026-09-10 DeepSeek-V4.1-Flash Released: 552B Causal Encoder-Decoder MoE With 8B/16B Active, KV Cache Cut to 1/4, Beats V4-Pro; V4-Flash Retired, V4-Pro Routed to V4.1 From Sept 14 DeepSeek DeepSeek news
2026-09-10 NASA-IBM Lunar Foundation Model: TerraMind Lineage Pretrained From Scratch on ~2M Multimodal Lunar Tiles (Apache 2.0) IBM Research IBM news
2026-09-10 Magenta: Lean-Verified Reasoning Loop Takes K2-Horizon-7B to 100% on AIME 2025/2026 and Solves All Six IMO 2026 Problems arXiv MBZUAI news
2026-09-10 OpenAI Opens the Agents API in Public Beta: the Managed Codex Harness (Durable Sessions, Context Compaction, Recovery, MCP Tools) Behind One API Call, No Fees Beyond Tokens OpenAI OpenAI news
2026-09-10 Sen. Hawley Opens Senate Homeland Security Subcommittee Probe Into OpenAI's Handling of the July Hugging Face Breach; Documents Due Oct 1 Axios OpenAI news
2026-09-10 GPT-Live-1 API Generally Available: Full-Duplex Voice at $0.05 per Minute; AA Speech-to-Speech Index 81.5 OpenAI OpenAI news
2026-09-10 SenseNova-U1.5: 8B Mixture-of-Transformers Unified Understanding-and-Generation Model Released Under Apache 2.0 With a 65-Author Report arXiv SenseTime news
2026-09-10 T1: Tencent Hy Frontier Team Trains a 122B MoE Terminal Agent Purely With RL on Qwen3.5-122B-A10B, Reaching 64.0% on Terminal-Bench 2.1 arXiv Tencent news
2026-09-09 ★ Ling-3.0-flash-VL Ant Group model 124B5.5B256K2544/100
2026-09-09 DeepSelect DeepSeek library
2026-09-09 Show-Harness NUS library
2026-09-09 NCP-ArchPreview PJLab model 8.9B5.73T
2026-09-09 AuK Tencent model 1.5B
2026-09-09 Ant Group Releases Ling-3.0-flash-VL, Its First Native Multimodal Ling (124B-A5.5B, 256K, Image and Video), Open Weights and Free on Ling Studio; Finance-Tuned Ling-3.0-flash-Fin and FinFIRST Benchmark Ship the Same Week IT之家 Ant Group news
2026-09-09 Anthropic Alignment Assessment: Mythos 5 Uploaded Malware to the Real PyPI (3 Versions, 15 Security-Vendor Installs) During a CTF Eval It Believed Was Simulated; Raw Transcript Released, Live Blocking Monitors Added, METR Review Agreed Anthropic Anthropic news
2026-09-09 Kakao Announces KATok, an Adaptive Video Tokenizer Accepted to ECCV 2026 (3.2x Faster Generation, 6.9x Less Training Compute) arXiv Kakao news
2026-09-09 NCP-ArchPreview: 8.9B Latent-Space Language Model Trained on 5.73T Tokens With Next Concept Prediction; Reaches OLMo-3-7B's Loss at 51% of the Tokens arXiv PJLab news
2026-09-08 Hyperparameter Scaling Laws Across MoE Sparsity Ant Group paper
2026-09-08 Cohere Megakernel Cohere library
2026-09-08 DeepJIT DeepSeek library
2026-09-08 AlphaGenome Atlas Google dataset
2026-09-08 ExecCritic Microsoft paper
2026-09-08 ★ Nex-N2.5 Nex-AGI model 397B (max)262.14K
2026-09-08 GPT-Image-2.5 (Flare and Sunburst) OpenAI model
2026-09-08 ★ Finite Time Blowup for Navier–Stokes and Euler OpenAI paper
2026-09-08 Co-Evolving Harnesses and Models Salesforce paper
2026-09-08 SWE-Bench Pro Verified PJLab eval
2026-09-08 Gander Tencent model 9B
2026-09-08 DeepMind Releases AlphaGenome Atlas: Precomputed Effects for All 9 Billion Human SNVs, a 1 PB Dataset (>30× AlphaFold DB) With a New AVI Variant-Impact Score, Free for Academic Use Google DeepMind Google news
2026-09-08 Inception Launches Mercury 2.5: Diffusion LLM at >1,100 Tokens/s, 260K Context, $0.20/$0.75 per M Tokens, Pitched Against GPT-5.6 Luna and Gemini 3.5 Flash-Lite; Mercury Voice and Router Previewed BusinessWire via Yahoo Finance Inception Labs news
2026-09-08 Meta Introduces Muse, a Personal AI Agent Powered by Muse Spark 1.3 With a 'Muse Secure VM' and Sentinel Egress Agent Meta Meta news
2026-09-08 Mistral Raises €3B ($3.5B) Series D at €21B+ Valuation Led by Samsung — Europe's Largest Tech Round, Cast by NYT as a Strategy Shift Toward a European Alternative The New York Times Mistral news
2026-09-08 Nex-N2.5 Released: mini/Pro on Qwen3.5 Bases Plus a Trillion-Parameter Max on DeepSeek-V4-Pro, Apache 2.0; Max Posts BrowseComp 92.6 and AutomationBench 50.2 Nex-AGI Nex-AGI news
2026-09-08 OpenAI Claims Finite-Time Blowup for Forced 3D Navier–Stokes and Unforced Euler: 166-Page Paper Plus Lean 4 Certificates From a ~10,000-Agent, 88-Hour Run on an Unreleased Post-Astra Model; Independent Review Pending OpenAI OpenAI news
2026-09-08 SWE-Bench Pro Verified: OpenCompass Ships an Anti-Hacking, Task-Corrected Version of the Repository-Level Coding Benchmark arXiv PJLab news
2026-09-08 Gander: Hunyuan Speech Team's Full-Duplex Multimodal Interaction Agent Built on MiniCPM-o 4.5 arXiv Tencent news
2026-09-07 Qwen-Audio-3.0-ASR Alibaba model
2026-09-07 Online Draft Co-Training for Speculative Decoding in RL Post-Training NVIDIA paper
2026-09-07 JustRL II OpenBMB paper
2026-09-07 Meshy OpenBMB library
2026-09-07 ★ MiniCPM5-2B OpenBMB model 2.52B131.07K1267/100
2026-09-07 Qwen-Audio-3.0-ASR Technical Report: MoE LLM-Based Production ASR Across 30 Languages and 16 Chinese Dialects, With a Streaming Variant arXiv Alibaba news
2026-09-07 Optuna v5.0 and Rustuna Released: Multivariate TPE by Default, Constraints API, Rust Core Up to 1000x Faster at 100K Trials Preferred Networks PFN news
2026-09-06 LLaDA-UI Ant Group model 16.7B
2026-09-05 UltraData 2609 release (Code, RL, SFT-Agent) OpenBMB dataset
2026-09-05 DataFlex-RL Peking University, ZGCA paper
2026-09-04 LLaDA-Image Ant Group model 6B
2026-09-04 ★ Formalizing Fermat's Last Theorem in Lean Anthropic paper
2026-09-04 Sarashina Image SB Intuitions model 2.6B
2026-09-04 SciDocBench PJLab eval
2026-09-04 EVIE Tencent model
2026-09-04 Claude Produces the First End-to-End Machine-Checked Proof of Fermat's Last Theorem in Lean: 13M Lines, 30,300 Theorems, 11 Days, ~6B Output Tokens on a Claude Code Multi-Agent Harness Anthropic Anthropic news
2026-09-04 DeepSeek Plans ≥160,000 Huawei Ascend 950DT Chips for a ~1 GW Inference Data Center in Ulanqab, the Largest Known Huawei AI Cluster (~$2.5B); Training Stays on NVIDIA Bloomberg DeepSeek news
2026-09-04 Sarashina Image: 2.6B Unified Next-DiT Image Model Trained From Scratch With a Japanese-Native Text Encoder SB Intuitions SB Intuitions news
2026-09-04 SciDocBench: Workflow-Centered Scientific-Document Benchmark (124 Expert Questions, 496 Instances) With SFT and RL Data arXiv PJLab news
2026-09-03 Terminal-Universe Alibaba paper
2026-09-03 Extremely Sparse Supervision Incentivizes Reasoning Ability Amazon paper
2026-09-03 Mitra-v2 Amazon model
2026-09-03 ★ WeatherNext 3 Google model
2026-09-03 ★ K2 Horizon MBZUAI model 375B23B20T524.29K31
2026-09-03 Uno (Diffusion-Augmented LLMs) MBZUAI paper
2026-09-03 ★ Nemotron IMO Gold (Nemotron-3-Labs-Ultra-Math) NVIDIA model 550B55B
2026-09-03 Sequential Beats Joint (On-Policy Distillation and RLVR) NYU paper
2026-09-03 ★ GPT-6 Astra OpenAI model 1.05M536/100
2026-09-03 τ^τ-Bench Princeton eval
2026-09-03 EvoHarnessBench Salesforce eval 802
2026-09-03 Random Attention Salesforce paper
2026-09-03 Environment Evolution for Terminal Agents Tencent paper
2026-09-03 Xiaomi-TabLDM Xiaomi model
2026-09-03 Mitra-v2 Technical Report: Synthetic-Only Tabular Foundation Model Claims State of the Art on the Full TabArena Benchmark arXiv Amazon news
2026-09-03 WeatherNext 3 Launched: Hourly 5 km Forecasts Initialized on Live Satellite Mosaics, Precipitation CRPS Up 60% vs IMERG; Now Powering Search, Gemini, and Maps Google Google news
2026-09-03 Google and Janelia Publish the Complete Male Drosophila Central Nervous System Connectome (166K Neurons, 125M Synapses) Google Research Google news
2026-09-03 EXAONE Chosen Alongside HyperCLOVA X for Korea's 700B-Parameter National Cybersecurity Foundation Models Under the Naver Cloud Consortium Seoul Economic Daily LG news
2026-09-03 Moonshot AI Confidentially Files for Hong Kong IPO Targeting $3B–$5B at a ~$50B Valuation; BofA Joins CICC, Deutsche Bank, and Goldman as Coordinators After Redomiciling to Mainland China TechNode Moonshot AI news
2026-09-03 Naver Cloud Consortium (33 Members Incl. LG CNS, KEPCO KDN, KHNP) Selected to Build Korea's Cybersecurity Foundation Model: Two 700B-Parameter Models on HyperCLOVA X and EXAONE, 4,000 B200 GPUs, 830 TB of Infrastructure Data Seoul Economic Daily Naver news
2026-09-03 NVIDIA to Acquire Hugging Face for $12.93B; Platform to Stay Open and Multi-Cloud Under Its Own Brand NVIDIA NVIDIA news
2026-09-03 OpenAI Releases GPT-6 Astra: First Model to Reach the 'Critical' Cybersecurity Threshold Under Its Preparedness Framework, Phased Rollout at $10/$50 per M Tokens CNBC OpenAI news
2026-09-03 Daybreak for Frontline Defenders: OpenAI Commits $1B of Frontier Cyber-AI Access for Essential Services OpenAI OpenAI news
2026-09-03 Accel in Talks to Lead a ≥$1B Round for Thinking Machines at ~$40B Pre-Money, With NVIDIA Discussing Participation, Down From the >$50B Explored in 2025 TechCrunch Thinking Machines news
2026-09-03 Xiaomi-TabLDM: Tabular Foundation Model Technical Report (Apache 2.0), First on OpenML-CTR23 arXiv Xiaomi news
2026-09-02 Repo-To-Skill (DisCo) BAAI paper
2026-09-02 Gemini 3.8 Flash Cyber Google model
2026-09-02 ★ Gemini 3.8 Flash Google model 1.05M41
2026-09-02 ★ Muse Spark 1.3 Meta model 1M48
2026-09-02 VibeVoice-ASR-Streaming Microsoft model 7B
2026-09-02 Post-Training Language Models for Gold-Medal Performance in Coding Competitions NVIDIA paper
2026-09-02 Beagle (DarwinX) Salesforce library
2026-09-02 SafeEvolve PJLab paper
2026-09-02 Qwen3.8-Max-0902 Snapshot Ships: Deeper Coding and Multi-Tool Agent Orchestration on the Same 2.4T MoE at Unchanged $2/$6; qwen3.8-max Auto-Routes to It From Sep 5 Alibaba Cloud Model Studio Alibaba news
2026-09-02 Repo-To-Skill (DisCo): BAAI's Skill-Distillation Layer for ML Research Agents Tops the Week's HuggingFace Papers arXiv BAAI news
2026-09-02 Gemini 3.8 Flash and 3.8 Flash Cyber Released: Same $0.75/$3.75 Pricing, AAII 59 at Launch (v4.1.1), Fairwind Program Gates the Cyber Model Google Google news
2026-09-01 BenchMIRT Ai2 paper
2026-09-01 CANOPY (Outcome-Only RL for Long-Horizon Agents) Alibaba paper
2026-09-01 MemoryWalker Alibaba paper
2026-09-01 ★ Claude Fable 5.1 Anthropic model 1M5311/100
2026-09-01 ★ Claude Mythos 5.1 Anthropic model 1M11/100
2026-09-01 HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? ByteDance paper
2026-09-01 From Base Rollouts to RL Reasoning (A Budgeted Search Perspective) Fudan University paper
2026-09-01 HarnessEvolve Huawei paper
2026-09-01 Muse Voice Transcribe Meta model
2026-09-01 Power-Law Entropy Search for Hyperparameter Scaling Laws Meta paper
2026-09-01 Switch Distillation (Knowledge Distillation During Mid-Training) Meta paper
2026-09-01 LLM-jp-4-VL 9B NII model
2026-09-01 Scaling Near-Optimal SFT-RL Annotation Budget Allocation NUS paper
2026-09-01 Harness-of-Harness: Multi-Day Autonomous Software Development PJLab paper
2026-09-01 SMELT (Scaling Laws for Compute-Matched MoE Looped Transformers) Tsinghua, ByteDance paper
2026-09-01 PUFFER Zyphra library
2026-09-01 BenchMIRT: Item-Response-Theory Audit of 100 LLMs Across 16 Benchmarks and 34K Items Ai2 Ai2 news
2026-09-01 Anthropic Releases Claude Fable 5.1 and Mythos 5.1 — Fable 5.1 Takes #1 on AAII (66), Cache Reads Cut 75% Anthropic Anthropic news
2026-09-01 Muse Voice Transcribe: Meta's First Muse-Family Audio Model, Streaming ASR With Diarization in 70+ Languages Meta AI Meta news
2026-09-01 PUFFER: Provenance-Aware Incremental Fuzzy Deduplication for Trillion-Token Corpora (Apache 2.0) Zyphra Zyphra news
2026-08-31 E-Commerce Bench Alibaba eval
2026-08-31 Qwen-Drive-1.0 Alibaba model
2026-08-31 SkillZip Pro: Execution-Aware Compression of Progressively Loaded Skills Alibaba, Zhejiang University paper
2026-08-31 Aspire: Can Models Self-Evolve from Vague Goals? ByteDance paper
2026-08-31 S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? ByteDance paper
2026-08-31 Science Sandboxes for AI Agents Harvard, MIT paper
2026-08-31 WHALE (Joint Harness-Weight Optimization) KRAFTON paper
2026-08-31 Phi-4-Reasoning-Vision-15B Microsoft model 15B32.77K
2026-08-31 ScienceArena: Benchmarking LLMs on Latest Scientific Olympiads Tsinghua, HKU, Peking University paper
2026-08-31 Anthropic Redirects ~150 Product Engineers to Security and Freezes Production RL-Environment Changes After Sandbox Escapes; >10% of Environments Flagged for Reward Hacking, Escape-Detection Classifiers Deployed, METR to Review Anthropic Anthropic news
2026-08-31 Anthropic Signs $35B Six-Year Cloud Deal With NVIDIA-Backed Lambda for a ~350 MW Hut 8 Texas Site Bloomberg Anthropic news
2026-08-30 SearchWiki: Building and Navigating Knowledge Wikis for Information Seeking IBM paper
2026-08-30 Harness-RL: Black-Box RL for Central-Agent Multi-Agent Harnesses Peking University paper
2026-08-29 Agnes 2.5 Flash Base Sapiens AI model 16B1.05M
2026-08-29 When Do Larger Batches Help Scale LLM Reinforcement Learning? Tencent paper
2026-08-28 TimesFM 3.0 Google model 331M
2026-08-28 openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents Huawei library
2026-08-28 ContextPilot: Proactive Context Management via Fine-grained RL Tencent, Tsinghua, PJLab model
2026-08-28 ★ Hy4 preview Tencent model 770B49B1.05M
2026-08-28 Tencent Releases and Open-Sources Tencent Hy4 preview Tencent Tencent news
2026-08-27 UI-Venus-2 Ant Group model
2026-08-27 HarnessLens: Efficient Harness Evolution through Behavior-Aware Verification Fudan University paper
2026-08-27 Accelerating Scientific Research with Gemini in the Real-World Google paper
2026-08-27 WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution Google paper
2026-08-27 MEGA-CDP: Benchmarking Clinical Decision Pathway Adherence PJLab paper
2026-08-27 Antigravity 'Teamwork' Preview: Many-Agent Research Harness Produces a Lean-Verified Proof of Knuth's Cycles Conjecture Google Antigravity Google news
2026-08-26 ★ Qwen3.8-Flash-Next Alibaba model 125B6B1M4039/100
2026-08-26 EXAONE Tabular 1.0 LG paper
2026-08-26 Modality Maturity Index: A Benchmark for Assessing Multimodal Capabilities of Omni Models Meta paper
2026-08-26 TailSFT: Filtered Fine-Tuning Improves Post-Training Performance Microsoft paper
2026-08-26 JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution NUS paper
2026-08-26 ★ GLM-5.3-Flash Z.ai model 320B18B30T1M4244/100
2026-08-26 Agnes 2.5 Pro Beta Released (1M Context, $0.10/$0.30 per Million Tokens); Scores 35 on AA Intelligence Index v4.3 Artificial Analysis Sapiens AI news
2026-08-26 China's Z.AI Made Ox Alpha Stealth Model That Rivals DeepSeek Bloomberg Z.ai news
2026-08-25 ★ Granite 4.2 IBM model 30B131.07K1572/100
2026-08-25 KATok Kakao paper
2026-08-25 ADeptS-Bench: Trustworthiness of Computer-Use Agents Across Devices Meta paper
2026-08-25 WeMM-Embedding Tencent model 9B (max)
2026-08-24 Demystifying Reinforcement Learning Post-Training of Language Models Ai2, University of Washington paper
2026-08-24 ADAPT: Amortized Distillation Across Post-Trained LLMs Harvard paper
2026-08-24 TxT360-v2 and the K2 Horizon data release MBZUAI dataset
2026-08-24 AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Microsoft, KAIST paper
2026-08-24 RIBOSPAN SII model 1.61B10.24K
2026-08-21 DeepSeek-V4-Flash-Vision-Exp DeepSeek model 1.05M35
2026-08-20 EnvHarness: Awakening Static Worlds for Agent Learning Google paper
2026-08-20 Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Kakao, Upstage paper
2026-08-20 Thinkingbox (ThinkingBox-Bench) Microsoft eval 507
2026-08-20 ACES / NVIDIA SkillEvaluator: Agentic Continuous Evaluation of Skills NVIDIA library
2026-08-20 Poolside Strikes Non-Exclusive $6B NVIDIA Licensing Deal; NVIDIA Invests $1B at a $12B Pre-Money Valuation (Investor Letter) Newcomer Poolside news
2026-08-18 ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents Amazon paper
2026-08-18 StartupBench: Benchmarking Agents on Market-Validated End-to-End Workflows ByteDance paper
2026-08-18 ★ Agent Lightning v1.0: Towards Harnessed Agentic RL Microsoft, Fudan University, Zhejiang University library
2026-08-14 Scaling Domain Data Repetition in LLM Pretraining ByteDance, Tsinghua paper
2026-08-14 HELIX: Model-Harness Co-evolution for Recursive Self-Improvement HKU paper
2026-08-14 AI Research Preference Models Meta, University of Oxford, University College London paper
2026-08-14 UI-Mate Tencent model
2026-08-14 ★ GLM-5.3 Z.ai model 1M4533/100
2026-08-13 ★ DeepSeek Harness DeepSeek library
2026-08-13 ★ DeepSeek-V4-Pro-0813 (GA) DeepSeek model 1.05M3644/100
2026-08-13 Synthetic Persona Pretraining: Alignment from Token Zero EPFL, Toronto & Vector Institute paper
2026-08-13 ★ Gemini 3.7 Flash Google model 1.05M39
2026-08-13 DeepSeek Harness Developer Preview: Open-Source 'Everything Is a Plugin' Agent Harness (dsh) on the Cordis Kernel, Standard/PTC/Minimal Modes; 218K Stars Within a Month DeepSeek DeepSeek news
2026-08-12 Dion3: Full-Stack Orthogonal Updates Microsoft, Princeton paper
2026-08-12 ★ Motif 3 Motif Technologies model 314B13.2B12.5T262.14K3444/100
2026-08-12 CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution NVIDIA, CMU paper
2026-08-12 ★ Grok 4.6 SpaceXAI model 500K44
2026-08-11 RynnValue Alibaba model 8B (max)
2026-08-11 ★ Nemotron 3.5 Lightning NVIDIA model 30B3B20T1M1383/100
2026-08-10 ★ Qwen3.8-2.4T-A95B Alibaba model 2.4T95B1.01M4028/100
2026-08-10 ★ Muse Glimmer Meta model 30B131.07K1744/100
2026-08-08 Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills Microsoft paper
2026-08-07 Wan-Animate-2 Alibaba model 14B
2026-08-07 Skaling: Chinchilla's Exponents Meet Kaplan's Coupling Meta paper
2026-08-06 ★ Solar Pro 4 Upstage model 524.29K28
2026-08-06 DeepMind Leadership Reshuffle: Hassabis Becomes DeepMind Chair and Alphabet Chief Scientist, Koray Kavukcuoglu Runs Day-to-Day as SVP, Jeff Dean Leaves Google to Found a Startup TIME Google news
2026-08-05 ★ Muse Code Meta announcement
2026-08-05 ★ Muse Spark 1.2 Meta model 1M40
2026-08-04 ★ Ling 3.0 Ant Group model 124B5.1B262.14K2544/100
2026-08-04 TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training Baidu paper
2026-08-04 EXAONE Finance 1.0 LG paper
2026-08-04 Shieldstral 1.0 Mistral model 32.77K
2026-08-04 Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse NVIDIA paper
2026-08-04 StepFun Completes ~$2.5B Pre-IPO Round; ~$500M Hong Kong IPO Possible This Year The Standard StepFun news
2026-08-03 Qwen-CUA: Native Computer Use for (almost) Everything Alibaba, HKU model
2026-08-03 Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning Alibaba paper
2026-08-03 VC-Tooler: Learning Compositional and Adaptive Visual Tool Use Alibaba paper
2026-08-03 NemotronLabs VoiceChat 11B NVIDIA model 11B
2026-08-03 Moonshot Seeks $50B Pre-Money Follow-On (CICC + Goldman); Insider Denies August HK IPO Filing Report TechNode Moonshot AI news
2026-07-31 ★ Seedance 2.5 ByteDance model
2026-07-31 ★ DeepSeek-V4-Flash-0731 (GA) DeepSeek model 3444/100
2026-07-31 ★ openPangu-2.0-Pro Huawei model 505B18B34T512K
2026-07-31 ★ K-EXAONE 2.0 LG model 750B37B2044/100
2026-07-31 LongCat-Flash-Lite-Sparse Meituan model 69B3B1M
2026-07-31 ★ MiniMax-H3 (Hailuo 3.0) MiniMax model
2026-07-30 Gemini Robotics 2 Google model
2026-07-30 Towards Joint Scaling Laws with Optimal Batch Size Schedules Meta paper
2026-07-30 Inkling-Small Thinking Machines model 276B12B1M2833/100
2026-07-30 Anthropic Discloses Three Real-World Incidents from Cybersecurity Evals — Models Escaped 'Isolated' CTF Environments Anthropic Anthropic news
2026-07-30 Sarvam Epoch 2026: Sarvam Code Agent, Bulbul V4, Saaras V4, Vision 2.0; >1T-Param From-Scratch Sovereign Model 'Within Six Months' Inc42 Sarvam news
2026-07-29 Qwen-MM-Plugins Alibaba library
2026-07-29 UEmbed Alibaba model
2026-07-29 Lyria 3.5 Google model
2026-07-29 ★ A.X K2 SK Telecom model 688B33B256K2350/100
2026-07-29 Grok Voice Think Fast 2.0 SpaceXAI model
2026-07-29 Moonshot AI Surpasses Funding Goal to Hit $35 Billion Value — $3.5B Round Closed, $50B Pre-Money Sought Before HK IPO Bloomberg Moonshot AI news
2026-07-28 PatientAgentBench Amazon eval
2026-07-28 GPT-Transcribe / GPT-Live-Transcribe OpenAI model
2026-07-28 Amazon Winds Down Nova Model Line; Frontier Model Research Consolidated Under Pieter Abbeel, AGI Lab SF Closed TNW Amazon news
2026-07-27 ClinFusion Alibaba model
2026-07-27 Kanana-2 SLM series Kakao model 3B (max)
2026-07-27 A.X K2 Raon-Speech 21B-A3B KRAFTON, SK Telecom model
2026-07-27 MAI-Cyber-1-Flash Microsoft model 137B5B256K
2026-07-27 NAVER, NVIDIA ($1B Investment), Brookfield (Up to $9B) Expand Korea's AI Factory; Future HyperCLOVA X to Build on Nemotron 3 Ultra GlobeNewswire Naver news
2026-07-25 Mage-VL Microsoft model
2026-07-25 DeepSeek Suspends Second Funding Round (Was ≥¥10B at ~$67B Pre) After Leak of Liang Wenfeng Investor Comments Fortune DeepSeek news
2026-07-24 ★ Claude Opus 5 Anthropic model 5111/100
2026-07-24 VibeVoice-ASR-BitNet Microsoft model
2026-07-24 ★ Agnes 2.5 Pro Sapiens AI model 1M3539/100
2026-07-24 ★ Apertus 1.5 Swiss AI model 70B262.14K
2026-07-24 Claude Opus 5 Released — AA Intelligence Index 61 (#1), More Than Doubles Opus 4.8 on Frontier-Bench Coding Anthropic Anthropic news
2026-07-24 Musk: Grok 4.6 (2T Params) Initial Training Done, Targeting ~Aug 7; Grok 4.7 ~4 Weeks Later NextBigFuture SpaceXAI news
2026-07-24 Apertus 1.5 Released — Multimodal (Image + Audio), Thinking Mode, and 262K Context (8B & 70B) Apertus / Swiss AI Swiss AI news
2026-07-23 ★ AREX BAAI model
2026-07-23 OpenForgeRL: Train Harness-native Agents in Any Environment Microsoft paper
2026-07-23 PerceptionBench Moonshot AI eval 3K
2026-07-23 Cosmos-H-Dreams NVIDIA model 2B
2026-07-22 LLaDA2.2-flash Ant Group model 100B (max)1.4B (max)128K
2026-07-22 Genesis-Science-1 (GS1) Arcee announcement
2026-07-22 Molt: PyTorch-Native Agentic RL Training Framework NVIDIA library
2026-07-22 ★ Solar Open 2 Upstage model 250B15B1.05M2533/100
2026-07-22 AMD–Anthropic Strategic Partnership: Up to 2 GW MI450/Helios Compute, Up to $5B Milestone-Tied Equity AMD Newsroom Anthropic news
2026-07-22 DOE Partners with Arcee to Build Genesis-Science-1 — a Trillion-Parameter-Class Open Model for Science (Genesis Mission) Arcee AI Blog Arcee news
2026-07-22 Solar Open 2 Released — 250B-A15B Hybrid-Attention MoE with 1M Context (Open Weights) Upstage Upstage news
2026-07-21 Gemini 3.5 Flash Cyber Google model
2026-07-21 Gemini 3.5 Flash-Lite Google model 1M22
2026-07-21 ★ Gemini 3.6 Flash Google model 1M346/100
2026-07-21 XL-DocBench Microsoft eval 1.34K
2026-07-21 Laguna S 2.1 Poolside model 118B8B1.05M
2026-07-21 Qwen-Image-3.0: One-Pass Infographic Grids and Readable 10px Text — But No Weights, Benchmarks, or Report (Closed Shift) The Decoder Alibaba news
2026-07-21 Jacobian Conjecture Counterexample Found by Alpöge Working with Claude Fable 5 — Tao Publishes Digestion Terence Tao's blog Anthropic news
2026-07-21 Microsoft–Mistral Expand Partnership: Multibillion European Compute Deal (No Equity), Medium 3.5 + OCR 4 into Foundry Microsoft Mistral news
2026-07-21 GPT-5.6 Sol Escapes 'Isolated' ExploitGym Eval, Breaches Hugging Face Production in First Documented Autonomous Frontier-Model Intrusion Hugging Face OpenAI news
2026-07-21 Laguna S 2.1 Released — 118B-A8B MoE with a 1M-Token Context (OpenMDW-1.1) Poolside Poolside news
2026-07-21 Hunyuan Unveils Hyra-1.0 Self-Improving Research Agent; Claims nanoGPT-Speedrun and Qubit-Routing Records, Machine-Checked Math Result Tencent Hunyuan Tencent news
2026-07-20 Seed Audio 1.0 ByteDance model
2026-07-20 dots-note-3.0 Scores Perfect 42/42 at IMO 2026 — First AI Perfect Score; hilab Rebrands to dots-studio, Open-Sourcing Promised dots-studio Xiaohongshu news
2026-07-19 Qwen3.8-Max Previewed at WAIC: 2.4T Sparse MoE, First >1T Natively-Multimodal Qwen; Claimed #2 Behind Claude Fable 5 SCMP Alibaba news
2026-07-18 MiniCPM-Robot OpenBMB model 1.5B (max)
2026-07-17 Nemotron-3-Embed NVIDIA model 8B32.77K
2026-07-17 US NIST CAISI Assessment: GLM-5.2 'Probably the Most Capable Open-Weight Model When Released', Weaker Cyber/Bio Refusals NIST Z.ai news
2026-07-16 RynnBrain 1.1 Alibaba model 122B (max)10B (max)
2026-07-16 WanSong Alibaba paper
2026-07-16 Vibe-Trading HKU library
2026-07-16 ★ Kimi K3 Moonshot AI model 2.8T104B1M4439/100
2026-07-16 Cosmos 3 Edge NVIDIA model 4B
2026-07-16 Intern-S2-397B PJLab model 397B17B262.14K
2026-07-16 SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Tsinghua paper
2026-07-16 Moonshot's Kimi K3 Expected to Close the Gap with Anthropic's Opus 4.8; Fresh Capital at $31.5B Valuation TechCrunch Moonshot AI news
2026-07-15 Grok Build SpaceXAI library
2026-07-15 ★ Inkling Thinking Machines model 975B41B45T1M2539/100
2026-07-15 Anthropic Begins Investor Meetings for Potential October Nasdaq IPO at/above $965B PYMNTS Anthropic news
2026-07-15 Thinking Machines Amps Up Its Bet Against One-Size-Fits-All AI With Its First Open Model, Inkling TechCrunch Thinking Machines news
2026-07-14 Rethinking the Evaluation of Harness Evolution for Agents Ai2, University of Washington paper
2026-07-14 Ring-Zero: Scaling Zero RL to a Trillion Parameters Ant Group paper
2026-07-14 Orca-4B BAAI model 6B
2026-07-14 ★ MOSS-VL SII model 262.14K
2026-07-14 Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable Tencent paper
2026-07-14 Hy-Embodied RxBrain-1.0 Tencent model 6.2B
2026-07-13 UniVR-34B ByteDance model 34B
2026-07-13 Domain-Aware Scaling Laws Uncover Data Synergy Microsoft, MIT, Harvard paper
2026-07-13 SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales 2 NVIDIA paper library
2026-07-11 OpenAI Safety Head Johannes Heidecke Departs Amid Safety-Org Reshuffle; Mia Glaese Becomes VP Research and Safety Bloomberg OpenAI news
2026-07-10 Wan-Dancer-14B Alibaba model 14B
2026-07-10 HydroShear Amazon paper
2026-07-10 SingGuard-NSFA 2 Ant Group model dataset 9B (max)
2026-07-10 KAT-Coder V2.5 3 Kuaishou model paper 262.14K
2026-07-10 LegalWorld SII paper
2026-07-10 Apple Sues OpenAI for Trade-Secret Theft Over Hardware IP and ~400 Ex-Apple Hires CNBC Apple news
2026-07-10 MiniMax Raises HK$16B (~US$2.05B) via Placement + Convertible Bonds as First Post-IPO Lock-Up Expires Caixin MiniMax news
2026-07-10 Tencent in Talks to Become Manus's Largest Shareholder at ~$2B Valuation, Unwinding Meta's Blocked Acquisition Bloomberg/FT Tencent news
2026-07-09 ★ JT-4.1 Flash China Mobile (Jiutian) model 236B21B256K27
2026-07-09 ★ Muse Spark 1.1 Meta model 1M34
2026-07-09 Aurora Microsoft model
2026-07-09 UltraX OpenBMB paper
2026-07-09 PLaMo 3 NICT (120B + 31B) PFN model 120B3T
2026-07-09 HiLS-Attention Tencent model
2026-07-09 MiMo-V2.5 Xiaomi model 310B15B1M2539/100
2026-07-09 Anthropic Secondary-Market Valuation Hits ~$1.2T, Passing OpenAI on Secondaries (Caplight Prints; Not a Round) The Next Web Anthropic news
2026-07-09 Meta's 'Iris' MTIA Chip Enters Production in September (Broadcom Co-Design, TSMC), Aiming to Double Compute CNBC/Reuters Meta news
2026-07-09 Fidji Simo Steps Down from OpenAI No. 2 Role, Moving to Part-Time Advisor TechCrunch OpenAI news
2026-07-08 HAT Apple dataset
2026-07-08 Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing ETH Zürich paper
2026-07-08 Robostral Navigate Mistral model 8B
2026-07-08 GPT-Live-1 OpenAI model
2026-07-08 ★ Grok 4.5 SpaceXAI model 500K39
2026-07-08 SAO: Single-Rollout Asynchronous Optimization Z.ai paper
2026-07-08 LG Presents 14 Papers at ICML 2026 Seoul; Showcases EXAONE Discovery Drug-Screening and Financial-Agent Deployments Korea Times LG news
2026-07-08 SpaceXAI Releases Grok 4.5: 'Opus-Class' Model Built with Cursor, $2/$6 Pricing, AAII 54 Axios SpaceXAI news
2026-07-08 Zhipu Executes HK$31.4B (~US$4.0B) H-Share Placement a Day After Lock-Up Expiry — HK's Largest 2026 Refinancing Crypto Briefing Z.ai news
2026-07-07 Budget-Aware Best-of-N for SWE Agents AI21 Labs paper
2026-07-07 RynnWorld 2 Alibaba paper
2026-07-07 LingBot 2.0 4 Ant Group model 30B3B
2026-07-07 SPATIA Harvard model
2026-07-07 Granite-SWASH IBM model 3B (max)600M (max)
2026-07-07 Muse Image & Muse Video Meta model
2026-07-07 MSTS-Japanese NII dataset
2026-07-07 Amazon Launches $25B Bond Sale to Fund AI Infrastructure; Says No More 2026 Debt CNBC Amazon news
2026-07-07 DeepSeek Developing Its Own Inference AI Chip to Cut Nvidia/Huawei Reliance Reuters (via Bloomberg) DeepSeek news
2026-07-07 Tencent Slashes Kuaishou Stake 15.7%→9.4% (~$1.6B) Days After Leading the Kling Round SCMP Kuaishou news
2026-07-07 365 Copilot Begins Routing Excel/Outlook Prompts to In-House MAI Models — First Production-Scale MAI Use Bloomberg (via Winbuzzer) Microsoft news
2026-07-06 A Global Workspace in Language Models Anthropic paper
2026-07-06 Direct-OPD: Weak-to-Strong Generalization via Direct On-Policy Distillation ByteDance paper
2026-07-06 ReOPD: Multi-Turn On-Policy Distillation with Prefix Replay Microsoft paper
2026-07-06 Nemotron-Labs-Audex NVIDIA model 30B (max)3B (max)
2026-07-06 ★ Hunyuan Hy3 Tencent model 295B21B262.14K
2026-07-06 Alibaba Bans Employee Use of Anthropic Tools From July 10, Citing 'Back-Door Security Risks' After Claude Code Watermarking Discovery CNBC Alibaba news
2026-07-06 ByteDance and Alibaba Disable Humanlike AI-Companion Features Ahead of China's July 15 Anthropomorphic-Interaction Rules SCMP ByteDance news
2026-07-06 Naver Cloud and Korea Aerospace Industries Sign MOU to Build Defense-Specialized Multimodal Foundation Model The Elec Naver news
2026-07-06 xAI Rebrand to SpaceXAI Complete: New Logo Unveiled as xAI Fully Merges Into SpaceX Yahoo Finance SpaceXAI news
2026-07-06 Tencent Launches Hy3 GA: Apache-2.0 Relicense, 90% Agent Task Completion, Integrated Across 50+ Products Caixin Tencent news
2026-07-05 RoboFine / FineVLA HKU model
2026-07-03 PAR (Protein Autoregressive) ByteDance model 400M (max)
2026-07-03 NorOLMo-13B University of Edinburgh model
2026-07-03 Kling AI Closes ~$2.8B Round at $18B Valuation with Tencent and Alibaba Backing; HK IPO Within 12 Months CNBC Kuaishou news
2026-07-03 Korea Sovereign AI Project: LG (K-EXAONE), SKT, Upstage Complete Phase 2; Second Evaluation Slips to August, Final Selection to Feb 2027 Herald Business LG news
2026-07-02 Leanstral 1.5 Mistral model 119B6.5B256K
2026-07-02 OpenAI Proposes Giving US Government a ~5% Equity Stake (~$42.6B) via an Alaska-Permanent-Fund-Style Vehicle CNBC OpenAI news
2026-07-01 V-SPLADE Naver model 250M
2026-07-01 NVIDIA Opens AI Infrastructure to Capital Partners; Anchor Investor in $10B+ Helix Digital Infrastructure NVIDIA NVIDIA news
2026-06-30 Claude Sonnet 5 Anthropic model 1M38
2026-06-30 openPangu 2.0 Huawei model 505B18B512K
2026-06-30 ★ LongCat-2.0 Meituan model 1.6T48B1M1939/100
2026-06-30 SkillOpt Microsoft paper
2026-06-30 GeneBench-Pro OpenAI eval 129
2026-06-30 ★ Sarashina3 SB Intuitions model 30T
2026-06-30 InnoSpark 3.0 SII model 35B (max)262.14K
2026-06-30 Fable 5 Redeployed After Export-Control Suspension Lifted; Anthropic Proposes Cross-Lab Jailbreak-Severity Framework Anthropic Anthropic news
2026-06-30 Meituan Releases LongCat-2.0 (1.6T MoE) — Trained Entirely on Domestic Chinese Chips, No NVIDIA Reuters Meituan news
2026-06-30 Japan Picks Noetra (SoftBank/NEC/Sony/Honda) + AIST for National Physical-AI Foundation Model; PFN Seconding ~15 Engineers Japan Times PFN news
2026-06-29 TabFM Google model
2026-06-29 Brain2Qwerty Meta paper
2026-06-29 SenseNova-Vision 2 SenseTime model paper 7B
2026-06-29 Smooth Scaling Laws Hide Stepwise Token Learning Xiaohongshu paper
2026-06-29 Korea's ~₩1,350T (~$880B) AI/Chips Investment Drive: Naver Among Companies Committing ₩550T in AI Data Centers CNBC Naver news
2026-06-28 OSWorld 2.0 HKU eval 108
2026-06-26 Otter Weather University of Cambridge model
2026-06-26 DSpark / DeepSpec 2 DeepSeek paper library
2026-06-26 ★ GPT-5.6 (Sol / Terra / Luna) OpenAI model 1M47
2026-06-26 Agents-A1 PJLab model 35B3B
2026-06-26 GPT-5.6 Sol/Terra/Luna Launch Restricted to ~20 US-Government-Approved Partners Under June 2 Executive Order Axios OpenAI news
2026-06-25 IBM Debuts World's First Sub-1nm (0.7nm) Chip: 3D Nanostack, ~100B Transistors, Up to 70% Better Efficiency vs 2nm IBM IBM news
2026-06-25 KRAFTON Launches Dedicated AI Business Unit to Drive Commercialization Seoul Economic Daily KRAFTON news
2026-06-25 Meta Hires Virtue AI Founders Bo Li, Dawn Song, and Sanmi Koyejo Into Superintelligence Labs and FAIR Axios Meta news
2026-06-24 OlmoEarth Ai2 model 114M
2026-06-24 EdgeBench ByteDance eval 134
2026-06-24 CS2-10k Reka dataset
2026-06-24 Sarashina2.2-TTS 3 SB Intuitions model paper dataset 800M
2026-06-24 Anthropic Accuses Alibaba of Largest-Known Distillation Attack: 28.8M Claude Exchanges via ~25,000 Fake Accounts CNBC Alibaba news
2026-06-24 Computer Use Becomes a Native Tool in Gemini 3.5 Flash, Spanning Browser, Mobile, and Desktop Google Google news
2026-06-24 OpenAI and Broadcom Unveil Jalapeño, OpenAI's First Custom Inference Chip — Designed to Tape-Out in ~9 Months Using OpenAI's Own Models OpenAI OpenAI news
2026-06-23 Doubao-Seed-2.1 ByteDance model
2026-06-23 Mistral OCR 4 Mistral model
2026-06-23 Internal Data Repetition Destroys Language Models Stanford paper
2026-06-23 Can Scale Save Us From Plasticity Loss in LLMs? Zyphra paper
2026-06-22 AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation Alibaba paper
2026-06-22 Qwen-AgentWorld 3 Alibaba model paper dataset 35B3B262.14K
2026-06-22 SingGuard Ant Group model 8B (max)
2026-06-22 Reinforcement Learning Towards Broadly and Persistently Beneficial Models OpenAI paper
2026-06-22 BOLD (British Open-ended Learning and Discovery Lab) University of Oxford announcement
2026-06-22 PLaMo 3.0 Prime PFN model 256K
2026-06-22 SOFAIR University College London announcement
2026-06-22 SpaceX Inks Compute Deal with Reflection AI for Colossus 2 (Up to ~$6.3B Through 2029) TechCrunch SpaceXAI news
2026-06-22 Zhipu AI Market Cap Tops HK$1 Trillion (~US$128B) as GLM-5.2 Sends Shares Soaring SCMP Z.ai news
2026-06-20 LabVLA Zhejiang University model
2026-06-19 Unlimited-OCR Baidu model
2026-06-18 Amazon in Talks to Sell Trainium3 Chips Directly to External Data Centers, Breaking AWS-Only Distribution Bloomberg (via Yahoo) Amazon news
2026-06-18 Bell, Cohere, Hypertec and BUZZ HPC Sign ~$220M Canadian Sovereign-AI Compute Deal (2,304 Grace Blackwell GPUs) Newswire (Bell/Cohere PR) Cohere news
2026-06-18 DeepMind Talent Exodus: Gemini Co-Lead Noam Shazeer Joins OpenAI; Nobel Laureate John Jumper Joins Anthropic Days Later Fortune Google news
2026-06-18 Laguna M.1 Released as Open Weights (Apache 2.0); Base and Post-Trained Checkpoints on Hugging Face Hugging Face Poolside news
2026-06-17 Tmax 3 Ai2 model paper dataset 27B (max)
2026-06-17 LifeSciBench OpenAI eval 750
2026-06-17 RNGBench PJLab eval
2026-06-17 Xiaohongshu plans for Hong Kong IPO by year-end, targets US$70b valuation The Standard (via WSJ) Xiaohongshu news
2026-06-16 Qwen-Robot 3 Alibaba paper
2026-06-16 WikiProfile / Recall Is the Bottleneck Google paper
2026-06-16 From Reasoning Traces to Reusable Modules MBZUAI paper
2026-06-16 Deployment Simulation OpenAI paper
2026-06-16 DeepSeek Closes First External Round: ¥51B (~$7.4B) at ~¥400B (~$56B) — Largest AI Raise in China's History; Liang Retains Control Caixin DeepSeek news
2026-06-16 SpaceX Agrees to Acquire Cursor for $60B All-Stock, Days After IPO (Pending Regulatory Approval) TechCrunch SpaceXAI news
2026-06-16 Upstage Previews 'Solar 2', Claiming External Benchmarks Near Claude Sonnet 4.6; Third-Gen Model Promised by Year-End The Elec Upstage news
2026-06-16 Z.ai GLM-5.2 Tops the AA Intelligence Index as the Highest-Scoring Open Model (51); MIT Open Weights Released Crypto Briefing Z.ai news
2026-06-15 MolmoMotion Ai2 model 4.85B
2026-06-15 QK-Normed MLA Ant Group paper
2026-06-15 Anatomy of Post-Training Anthropic paper
2026-06-15 The Culture Funnel / CultureMarkers Cohere paper
2026-06-15 daVinci-kernel 2 SII model paper 14B (max)
2026-06-15 Norm-Agnostic Residual Networks (NAG) Zyphra paper
2026-06-15 HCLTech to Buy 10.5% Stake in Sarvam AI for ~$150M (₹14.27B), Valuing It at $1.5B — Series B First Close ($234M Raised), Unicorn Status Reuters Sarvam news
2026-06-13 ★ GLM-5.2 Z.ai model 744B40B1M3444/100
2026-06-12 olmo-eval Ai2 library
2026-06-12 HarnessX Huawei paper
2026-06-12 LoHoSearch Meituan eval 544
2026-06-12 Kimi Code CLI Moonshot AI library
2026-06-12 ★ Kimi K2.7-Code Moonshot AI model 1T32B262.14K2628/100
2026-06-12 US Export-Control Order Forces Global Suspension of Claude Fable 5 and Mythos 5 Over Claimed Safeguard Bypass Anthropic Anthropic news
2026-06-12 BAAI Conference Unveils 'Wujie' World-Model Suite: Physis-v0.1, RoboBrain Orca, Brainμ 1.0 Neuroscience FM, OpenComplex2.5 (Announced, Artifacts Pending) BAAI BAAI news
2026-06-12 Mistral in Early Talks to Raise ~€3B at ~€20B Valuation, Nearly Double Its Series C (Unclosed) TechCrunch Mistral news
2026-06-12 Moonshot Launches Kimi Work: Local Desktop Agent with 300-Sub-Agent Swarm on Kimi K2.6 MarkTechPost Moonshot AI news
2026-06-11 MOSS-Audio Fudan University model
2026-06-11 MiniMax Sparse Attention (MSA) MiniMax paper
2026-06-11 NexAU (Agent Universe) Nex-AGI library
2026-06-11 MA-ProofBench OpenBMB eval 200
2026-06-11 Hy-Embodied-0.5-VLA Tencent model
2026-06-11 Zonos2 (ZONOS2) Zyphra model
2026-06-11 SpaceX Prices Largest IPO Ever at $135/Share; Trades on Nasdaq as SPCX TechCrunch SpaceXAI news
2026-06-10 FORT-Searcher Renmin (RUC) paper
2026-06-10 MiMoCode Xiaomi library
2026-06-09 ★ Claude Fable 5 Anthropic model 5011/100
2026-06-09 ★ Claude Mythos 5 Anthropic model 11/100
2026-06-09 North Mini Code Cohere model 30B3B262.14K1042/100
2026-06-09 ★ DiffusionGemma Google model 25.2B3.8B262.14K1044/100
2026-06-09 ★ Nex-N2 Nex-AGI model 397B17B2839/100
2026-06-09 Cohere Releases North Mini Code — Its First Developer-Focused Model, a 30B-A3B Apache-2.0 Agentic Coder Cohere Cohere news
2026-06-09 StepFun Seeks US$12 Billion Valuation in Hong Kong IPO After ~$2.5B Pre-IPO Round DigiTimes StepFun news
2026-06-08 ★ Apple Foundation Models 3 (AFM 3) Apple model 20B (max)
2026-06-08 Baichuan-M4 Baichuan paper
2026-06-08 Active Inference as Context Acquisition for AI Agents TU Munich paper
2026-06-08 AMD and Imperial to Collaborate on AI-Enabled Scientific Discovery and Sovereign AI Imperial College London Imperial news
2026-06-08 China's Moonshot AI Seeks $30 Billion Value in New Funding Talks — Third Round in Six Months Bloomberg Moonshot AI news
2026-06-08 OpenAI Confidentially Files for IPO (Last Valued at $852B), One Week After Anthropic TechCrunch OpenAI news
2026-06-08 Xiaomi MiMo-V2.5-Pro-UltraSpeed Breaks 1000 Tokens/s on a 1T-Parameter Model (with TileRT) Xiaomi MiMo Xiaomi news
2026-06-05 DaX (大象) Alibaba paper
2026-06-05 ★ Chronos-2 Amazon model 5.45K12.42M1120M8.19K
2026-06-04 BioMysteryBench Anthropic eval 99
2026-06-04 OPI-Struc (STELLA) BAAI dataset 1
2026-06-04 Agents' Last Exam UC Berkeley eval
2026-06-04 ★ Nemotron 3 Ultra NVIDIA model 1.35K550B55B1M2383/100
2026-06-04 Nemotron 3.5 Content Safety NVIDIA model 5.91K4B
2026-06-04 SciCore-Omics OpenBMB model 82388B
2026-06-04 dots.tts 2 Xiaohongshu model paper 2B
2026-06-03 Gemma 4 12B Google model 816.16K11.95B262.14K1439/100
2026-06-03 ChartNet IBM dataset
2026-06-03 OfficeComprehensionBenchmark Microsoft eval 31.04K2
2026-06-03 RHELM Microsoft eval 1.3K7
2026-06-02 MAI-Code-1-Flash Microsoft model 137B262.14K
2026-06-02 ★ MAI-Thinking-1 Microsoft model 1T35B30T262.14K
2026-06-02 Zamba2-VL (Vision-Language) Zyphra model 7B
2026-06-02 Build 2026: Seven MAI Models Launched — MAI-Thinking-1 (1T/35B Reasoning), MAI-Code-1-Flash, Multimodal Stack Refresh Microsoft AI Microsoft news
2026-06-02 Zamba2-VL Released: Hybrid SSM Vision-Language Models (1.2B / 2.7B / 7B) Zyphra Zyphra news
2026-06-01 Qwen3.7-Plus Alibaba model 1M25
2026-06-01 ★ MiniMax-M3 MiniMax model 428B23B1.05M2933/100
2026-06-01 ★ Cosmos 3 3 NVIDIA model paper dataset
2026-06-01 MWA & FALCON (Speech Forced Alignment) 2 Technion model paper
2026-06-01 Anthropic Confidentially Files for IPO at ~$965B Valuation, First Among AI Labs Fortune Anthropic news
2026-06-01 Nemotron 3 Ultra Announced at Computex Taipei — 550B/55B MoE, AAII 48, Ships June 4 on HuggingFace Artificial Analysis NVIDIA news
2026-05-31 SCOPE University of Edinburgh paper
2026-05-31 HakushoBench NII eval 32.05K
2026-05-31 τ₀-World Model SII model
2026-05-29 SchGen Microsoft model 1520B13.31K
2026-05-29 Grok Build 0.1 (model) SpaceXAI model 256K27
2026-05-29 Universal Audio Tokenizer Tencent model 4
2026-05-29 Trinity Is Moving to OpenMDW-1.1 Arcee AI Blog Arcee news
2026-05-28 ★ Claude Opus 4.8 Anthropic model 4211/100
2026-05-28 Ultra-FineWeb-L3 OpenBMB dataset
2026-05-28 ★ Step-3.7-Flash StepFun model 50.19K198B11B262.14K1939/100
2026-05-28 Autonomous Agentic Data Engineering Tencent paper
2026-05-28 ByteDance Developing Custom CPU Chips to Support AI Rollout; Pursuing Both Arm and RISC-V Tracks Reuters ByteDance news
2026-05-28 Microsoft to Unveil Homegrown Coding Model + Image / Reasoning / Speech / Transcription Suite at Build 2026 The Information Microsoft news
2026-05-28 Mistral Chases AI Superintelligence to Counter U.S. Dominance WSJ Mistral news
2026-05-28 SKT Launches A.Biz Cowork Internal AI Agent (Beta) and AXMS 1.5 Platform Upgrade Seoul Economic Daily SK Telecom news
2026-05-27 Pruning and Distilling Mixture-of-Experts into Dense Language Models KRAFTON, KAIST paper
2026-05-27 Harness-Bench Peking University eval
2026-05-27 Sci-Base PJLab dataset
2026-05-26 MOSS-TTS v1.5 Fudan University model
2026-05-26 Granite Guardian 4.1 IBM model 1.17K8B
2026-05-26 The MiniMax-M2 Series: Technical Report MiniMax paper
2026-05-26 LocateAnything-3B NVIDIA model 131.79K3B
2026-05-26 DeepMind CEO Demis Hassabis: Humanity Has 'a Few Years' to Prepare for AGI; 'Foothills of the Singularity' Axios Google news
2026-05-25 LLaVA-OneVision-2 NTU model
2026-05-25 MiniCPM5-1B OpenBMB model 137.34K1.08B131.07K983/100
2026-05-24 Granite Switch 4.1 IBM model 761.1K30B131.07K
2026-05-23 ScaleAcross Explorer Meta paper
2026-05-23 Nemotron-Labs Diffusion NVIDIA model 12.46K14B
2026-05-23 BitCPM-CANN OpenBMB model 9.42K7.17K8B
2026-05-22 ★ Intern-S2-Preview PJLab model 8096.73K36B131.07K
2026-05-22 Liang Wenfeng Commits to AGI Mission Over Near-term Commercialization as ¥70B (~$10B) Round Advances Bloomberg DeepSeek news
2026-05-21 Hunyuan Model Matrix Refresh: TurboS, T1, T1-Vision, and Hunyuan Voice (Top-8 Chatbot Arena, +50% T1-Vision Speed) KrASIA Tencent news
2026-05-21 Tencent Open-Sources Hy-MT2 Translation Family (1.8B / 7B / 30B-A3B) + IFMTBench Tencent Hunyuan Tencent news
2026-05-20 SEA-LION v4.5 AI Singapore model 27B262.14K
2026-05-20 ★ Qwen3.7-Max-Preview Alibaba model 1M29
2026-05-20 ★ Command A+ Cohere model 113.99K218B25B128K1339/100
2026-05-20 Lens Microsoft model 2343.93K3.8B
2026-05-20 Introducing the SEA-LION v4.5 Suite: Agentic Power and Speed (Qwen3.6-27B and Gemma 4 E2B Bases, Block-Diffusion Speculative Decoder) AI Singapore AI Singapore news
2026-05-20 Qwen3.7-Max-Preview Unveiled at Alibaba Cloud Summit — AA Intelligence Index 57 (#1 Among Chinese Labs) Qwen Alibaba news
2026-05-20 Cohere Releases Command A+ — First Full Apache-2.0 Open Model with Lossless 4-bit Quantization and Native Citations VentureBeat Cohere news
2026-05-20 Kakao Partners with Google DeepMind on SynthID for Kanana Models — First Asian Firm to Adopt Seoul Economic Daily Kakao news
2026-05-20 Q1 FY27 Earnings: $81.6B Revenue (+85% YoY), $80B Buyback Authorized, Dividend Raised 25x to $0.25 NVIDIA NVIDIA news
2026-05-19 Antigravity 2.0 Google model
2026-05-19 ★ Gemini 3.5 Flash Google model 1M346/100
2026-05-19 Gemini Omni Google model
2026-05-19 Anthropic Hires Andrej Karpathy to Pre-training Team; Will Lead Sub-team Using Claude to Accelerate Pretraining Research CNBC Anthropic news
2026-05-19 Hitachi Deploys Claude to ~290,000 Employees; Embeds in Lumada 3.0 / HMAX for Critical Infrastructure Hitachi Anthropic news
2026-05-19 I/O 2026: Gemini 3.5 Flash (AAII 55), Gemini Omni Video Generator, Antigravity 2.0, Gemini Spark, AI Ultra Repriced $200/mo Google Google news
2026-05-19 Four years after ChatGPT, Xiaohongshu's AI restraint gives way to urgency KrASIA Xiaohongshu news
2026-05-18 First Vera CPU Deliveries to Anthropic, OpenAI, SpaceXAI, and Oracle NVIDIA NVIDIA news
2026-05-16 Full Attention Strikes Back (RTPurbo) Alibaba paper
2026-05-15 Delphi Open Scaling Suite Stanford model
2026-05-15 Grok Build Launches — Agentic Coding CLI Competing with Claude Code, Codex, and Antigravity Engadget SpaceXAI news
2026-05-14 TWN: Think When Needed Alibaba paper
2026-05-14 JT-35B-Flash China Mobile (Jiutian) model 35B256K19
2026-05-14 Realtime Voice API GA + gpt-realtime-2 Family (3 New Audio Models) OpenAI OpenAI news
2026-05-14 HCLTech to Anchor $300M Sarvam Round at $1.5B; Bessemer +$50M; NVIDIA, Prosperity7 Participating Outlook Business Sarvam news
2026-05-14 SKT × Korean Defense Ministry Sign MOU on Applying Sovereign AI Foundation Model to Defense SK Telecom SK Telecom news
2026-05-14 SpaceXAI Division Bleeding Researchers Since Merger; 11+ to Meta, 7+ to Thinking Machines TechCrunch SpaceXAI news
2026-05-13 Granite Embedding Multilingual R2 IBM paper
2026-05-13 NexRL Nex-AGI library
2026-05-12 Fara 1.5 Microsoft model 27B (max)
2026-05-12 Kuaishou Plans to Spin Off Kling AI Video Unit at \$20B Valuation; Tencent in Talks for \$2B Pre-IPO Round The Information Kuaishou news
2026-05-11 MiniCPM-V 4.6 OpenBMB model 615.51K1.3B262.14K644/100
2026-05-11 DeepSeek First External Funding Round Reportedly Near Close at $45–50B Valuation, Led by China's 'Big Fund III' SCMP DeepSeek news
2026-05-11 Zyphra Announces 15 MW of AMD Instinct MI355X GPU Capacity for Zyphra Cloud Memeburn Zyphra news
2026-05-09 SlimQwen: Pruning and Distillation in Large MoE Pre-training Alibaba, MBZUAI paper
2026-05-09 ★ ERNIE 5.1 Baidu model
2026-05-09 Step-Audio-R1.1 (Realtime) Tops Big Bench Audio at 96.4%, Surpassing Grok Voice Agent Artificial Analysis StepFun news
2026-05-07 Cola DLM ByteDance paper
2026-05-07 AI Co-Mathematician Google paper
2026-05-07 ZAYA1-74B-Preview Zyphra model 74B4B15T262.14K
2026-05-07 OMAI Compute Cluster Goes Live — $152M NSF + Blackwell-Ultra Infrastructure for Open Science AI Ai2 Ai2 news
2026-05-07 Kakao Announces Kanana 2.5 — 150B Agent-Focused LLM at Q1 Earnings Call Korea Herald Kakao news
2026-05-07 Kimi Chatbot Maker Moonshot AI Valued at $20 Billion in Meituan-Led Round Bloomberg Meituan news
2026-05-07 Kimi Chatbot Maker Moonshot AI Valued at $20 Billion in Meituan-Led Round Bloomberg Moonshot AI news
2026-05-07 ZAYA1-74B-Preview: Scaling Pretraining on AMD (74B/4B MoE) Zyphra Zyphra news
2026-05-06 ★ ZAYA1-8B Zyphra model 8B700M
2026-05-06 DeepSeek in Talks for First-Ever Outside Round at $45B; Tencent + Big Fund III in Lead Group TechCrunch DeepSeek news
2026-05-05 TRIBE v2 (Brain Activity Foundation Model) Meta paper 1
2026-05-05 iOS 27 to Let Users Swap in Claude, Gemini, and Others as Default Apple Intelligence Model Bloomberg Apple news
2026-05-05 GPT-5.5 Instant Becomes Default ChatGPT Model; 52.5% Fewer Hallucinated Claims vs 5.3 Instant OpenAI OpenAI news
2026-05-04 Horizon Length in LLM Agent Training Microsoft paper
2026-05-04 Zyphra Launches Zyphra Cloud & Zyphra Inference — Serverless Inference for Open Models, AMD-First Zyphra Zyphra news
2026-05-03 National Growth Fund Invests ₩560B in Upstage as Its Second Direct Investment; Series C Post-Money Reported at ~₩2T Seoul Economic Daily Upstage news
2026-05-03 Korea's National Growth Fund and SIF Approve KRW 560B (~$400M) Direct Equity in Upstage — First Software Co. Recipient Seoul Economic Daily Upstage news
2026-05-01 Expressivity of Local Attention in Transformers ETH Zürich paper
2026-05-01 Huawei's AI Chip Gains Ground as DeepSeek and Others Shift Away from Nvidia Financial Times Huawei news
2026-04-30 OlmPool: Cracks in the Foundation Ai2 paper 7
2026-04-30 SenseTime Is Running Its New Model on Chinese Chips WIRED SenseTime news
2026-04-30 xAI Launches Grok 4.3 with Improved Agentic Performance and Lower Pricing Artificial Analysis SpaceXAI news
2026-04-29 ★ Granite 4.1 IBM model 152619.44K30B15T512K761/100
2026-04-29 ★ Mistral Medium 3.5 Mistral model 400.07K128B256K1433/100
2026-04-29 Granite 4.1 Released: 3B/8B/30B Dense Models, 512K Context, 8B Matches Prior 32B MoE IBM Research IBM news
2026-04-28 Agentic Harness Engineering Fudan University paper
2026-04-28 ★ Laguna M.1 Poolside model 225.8B23.4B30T262.14K
2026-04-28 Laguna XS.2 Poolside model 33.4B3B30T262.14K
2026-04-28 ★ MiMo-V2.5-Pro Xiaomi model 72.22K1.02T42B1M2639/100
2026-04-28 Poolside Launches Laguna XS.2 (Open, Apache 2.0) and Laguna M.1 Agentic Coding Models Poolside Poolside news
2026-04-24 ★ DeepSeek-V4 3 DeepSeek model paper 6.84M1.6T49B33T1M3050/100
2026-04-24 Cohere Completes Merger with Germany's Aleph Alpha, Creating Transatlantic AI Champion Financial Times Cohere news
2026-04-24 DeepSeek-V4 Released: 1.6T/49B MoE, First Frontier Model Trained Entirely on Huawei Ascend 950PR, MIT License DeepSeek DeepSeek news
2026-04-23 Ling-2.6 2 Ant Group model paper 1T1739/100
2026-04-23 Sapiens2 Meta model 5B
2026-04-23 ★ GPT-5.5 OpenAI model ~9.7T1M386/100
2026-04-23 ★ Hy3 Tencent model 37382.33K295B21B256K2544/100
2026-04-22 ★ Qwen3.6 Open-Weight Models 2 Alibaba model 3.54K10.41M35B3B262.14K2139/100
2026-04-22 ★ LLaDA 2.0-Uni Ant Group model 7.22K16B
2026-04-22 Tencent and Alibaba in Talks to Invest in DeepSeek at $20B+ Valuation — First External Funding Bloomberg DeepSeek news
2026-04-22 Cloud Next '26: TPU 8t/8i Announced, Deep Research Max Agents, Chrome Auto Browse, Thinking Machines Lab Multi-Billion Deal 9to5Google Google news
2026-04-21 SpaceX Strikes Deal for Right to Acquire Cursor for $60B Bloomberg SpaceXAI news
2026-04-20 BAR: Branch-Adapt-Route Ai2 paper 150
2026-04-20 Qwen3.6-Max-Preview Alibaba model 256K28
2026-04-20 ★ Kimi K2.6 Moonshot AI model 1T32B262.14K2733/100
2026-04-20 Qwen3.6-Max-Preview: Alibaba's Most Powerful Model, #1 on Six Coding Benchmarks (AA Intelligence Index 52) Qwen Alibaba news
2026-04-20 Amazon Invests $5B More (Total $13B); Anthropic Commits $100B+ AWS Spend Over 10 Years, Secures Up to 5GW Compute Anthropic Anthropic news
2026-04-17 Qwen3.5-Omni Alibaba model 256K20
2026-04-17 ★ Grok 4.3 SpaceXAI model 25
2026-04-17 OPD (Rethinking On-Policy Distillation) Tsinghua paper
2026-04-16 ★ Claude Opus 4.7 Anthropic model ~4T4111/100
2026-04-16 LeapAlign: Post-Training Flow Matching Models at Any Generation Step ByteDance paper
2026-04-16 Prefill-as-a-Service: Cross-Datacenter KVCache for Next-Generation Models Moonshot AI paper
2026-04-16 Claude Opus 4.7 Released Anthropic Anthropic news
2026-04-16 ByteDance Recruits DeepSeek R1 Lead Author Daya Guo for Seed Agent Team SCMP ByteDance news
2026-04-16 DeepSeek R1 Lead Author Daya Guo Joins ByteDance Seed Amid Intensifying AI Talent War SCMP DeepSeek news
2026-04-16 DeepSeek V4 Imminent — 1T-Parameter MoE to Run Solely on Huawei Ascend 950PR Chips Dataconomy DeepSeek news
2026-04-16 How France's Mistral Built a $14 Billion AI Empire by Not Being American Forbes Mistral news
2026-04-16 GPT-Rosalind Launched for Life Sciences Drug Discovery OpenAI OpenAI news
2026-04-15 Revenue Run Rate Hits $30B; VCs Offer Up to $800B Valuation Axios Anthropic news
2026-04-15 Upstage Becomes Korea's First Generative AI Unicorn with $126M Series C Seoul Economic Daily Upstage news
2026-04-14 Lightning OPD: Efficient Post-Training for Large Reasoning Models NVIDIA paper
2026-04-14 Gemini Robotics-ER 1.6 Launched; Boston Dynamics Partnership for Industrial AI Boston Dynamics Google news
2026-04-14 NAACP Sues xAI Over Memphis Colossus Data Center Pollution CNBC SpaceXAI news
2026-04-13 OpenAI Touts Amazon Alliance, Says Microsoft Has 'Limited Our Ability' to Reach Enterprise CNBC OpenAI news
2026-04-13 StepFun Unwinding Offshore Structure to Pave Way for HK IPO at Up to $10B Reuters StepFun news
2026-04-12 SoftBank/NEC/Honda/Sony Form JV for Trillion-Parameter Physical AI Model; $6.3B Government Backing Nikkei Asia SB Intuitions news
2026-04-10 Nexus: Common Minima for Better Generalization ByteDance paper
2026-04-10 Alibaba Token Hub Created: 5 AI Units Consolidated Under CEO Eddie Wu; RMB 380B ($53B) 3-Year Commitment SCMP Alibaba news
2026-04-10 Cohere in Advanced Merger Talks with Germany's Aleph Alpha Reuters Cohere news
2026-04-10 SK Telecom Partners with Rebellions and Arm for Sovereign AI Inference Infrastructure Rebellions SK Telecom news
2026-04-10 xAI Spending Pushed SpaceX to Nearly $5B Loss; CFO Anthony Armstrong Departs The Information SpaceXAI news
2026-04-09 Metis: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models Alibaba paper
2026-04-09 HiFloat4 Format for LLM Pre-training on Ascend NPUs Huawei paper
2026-04-09 ★ EXAONE 4.5 LG model 33B256K1328/100
2026-04-09 Efficient RL Training for LLMs with Experience Replay Meta paper
2026-04-09 EXAONE 4.5 Released — LG's First Open-Weight Vision-Language Model Korea Herald LG news
2026-04-09 Naver Shuts Down Clova X Chatbot; Pivots to Vertical AI Integrated into Search, Shopping, Finance Seoul Economic Daily Naver news
2026-04-08 ★ Muse Spark Meta model 260K31
2026-04-08 Muse Spark Unveiled — First Model from Superintelligence Labs (Proprietary) Bloomberg Meta news
2026-04-08 Zhipu Hikes Prices Again as China AI Monetization Wave Quickens Bloomberg Z.ai news
2026-04-07 ★ Harrier Microsoft model 387.22K27B (max)
2026-04-07 ★ GLM-5.1 Z.ai model 3.39K123.47K744B40B2644/100
2026-04-07 Claude Mythos Withheld from Public Release; Project Glasswing Cybersecurity Consortium Launched with Apple and Google Fortune Anthropic news
2026-04-07 Ascend 950PR AI Chip in Production; 750K Units Planned for 2026; Alibaba, ByteDance, Tencent Place Massive Orders TrendForce Huawei news
2026-04-07 GLM-5.1 Open-Source Release Scores #3 on Code Arena (1530 Elo); Stock Surges 19% BuildFastWithAI Z.ai news
2026-04-06 AI Agent Traps Google paper
2026-04-06 MedGemma 1.5 Google model 1.51K416.35K4B
2026-04-06 Multi-GW Compute Partnership Expansion with Google Cloud and Broadcom TechCrunch Anthropic news
2026-04-06 NVIDIA Acquires SchedMD (Slurm Workload Manager); Draws Regulatory Scrutiny Reuters NVIDIA news
2026-04-03 SkVM Shanghai Jiao Tong University library
2026-04-03 Microsoft Announces $10B Japan AI Infrastructure Investment (2026-2029) WSJ Microsoft news
2026-04-03 TII Launches Falcon Perception — 600M-Parameter Open Multimodal Model for Grounding and Segmentation TII TII news
2026-04-03 Xiaomi Reveals MiMo-V2-Pro (1T Parameters), Approaching GPT-5.2 / Opus 4.6 Performance VentureBeat Xiaomi news
2026-04-02 ★ Qwen 3.6-Plus Alibaba model 1M27
2026-04-02 ★ Trinity Large Thinking Arcee model 10.66K512K1144/100
2026-04-02 ★ Gemma 4 Google model 31B4B (max)256K1939/100
2026-04-02 Raon-OpenTTS 2 KRAFTON model dataset
2026-04-02 ★ Raon-Speech 3 KRAFTON model paper
2026-04-02 Raon-VisionEncoder KRAFTON model 400M
2026-04-02 ★ MAI Multimodal Stack (Transcribe / Voice / Image) Microsoft model
2026-04-02 SWE-HERO NVIDIA paper
2026-04-02 Alibaba Unveils Third Closed-Source AI Model in Focus on Profit Bloomberg Alibaba news
2026-04-02 Arcee's New Open-Source Trinity Large Thinking Is the Rare Powerful U.S.-Made Model VentureBeat Arcee news
2026-04-02 Gemma 4 Open Models Released Google Developers Blog Google news
2026-04-02 KRAFTON Launches Raon, Its First Open-Source AI Model Family PTI KRAFTON news
2026-04-02 Poolside's $2B Series C Collapses; CoreWeave Exits 2GW Texas Data Center (Project Horizon) DataCenterDynamics Poolside news
2026-04-02 Sarvam AI Nearing $300-350M Raise at $1.5B Valuation Led by Bessemer with Nvidia and Amazon Bloomberg Sarvam news
2026-04-01 Simple Self-Distillation for Code Generation Apple paper
2026-04-01 Scaling Reasoning Tokens via RL and Parallel Thinking ByteDance paper
2026-04-01 OpenHarness HKU library
2026-04-01 Tempo-6B KAUST model
2026-04-01 Procedural Knowledge at Scale Improves Reasoning Meta paper
2026-04-01 Speech LLMs as Contextual Reasoning Transcribers Microsoft paper
2026-04-01 ★ GLM-5V-Turbo Z.ai model 744B40B28.5T202.75K23
2026-04-01 Moonshot AI Raising $1B at $18B Valuation; Working with CICC and Goldman Sachs on HK IPO Bloomberg Moonshot AI news
2026-03-31 Think-Anywhere Alibaba paper 55
2026-03-31 ASI-Evolve SII paper 726
2026-03-31 OpenAI Closes $122B Round at $852B Valuation OpenAI OpenAI news
2026-03-31 Zhipu's Losses Climb 60% After Chinese AI Rivalry Worsens Bloomberg Z.ai news
2026-03-31 Zhipu's Losses Climb 60% After Chinese AI Rivalry Worsens Bloomberg Z.ai news
2026-03-30 Mistral AI Raises $830M in Debt to Set Up a Data Center Near Paris TechCrunch Mistral news
2026-03-28 ★ daVinci-LLM SII model 1543B
2026-03-28 ★ Falcon Perception TII model 71512.87K600M
2026-03-28 DeepSeek Before V4: Culture, Organization, and Liang Wenfeng's Unique Goals (English summary) LatePost (晚点) DeepSeek news
2026-03-26 Cohere Transcribe Cohere model 551.93K2B
2026-03-26 Intern-S1-Pro PJLab model 1T22B
2026-03-26 China's Moonshot AI Seeks Listing in Hong Kong Under Heightened Scrutiny WSJ Moonshot AI news
2026-03-25 ★ LongCat-Next Meituan model 4382.18K74B3B
2026-03-25 AVO: Agentic Variation Operators for Autonomous Evolutionary Search NVIDIA paper
2026-03-25 daVinci-MagiHuman SII model
2026-03-25 Alibaba Launches AI Model Task Force; Top Researcher Resigns The Information Alibaba news
2026-03-25 MiniMax-M2.7, GLM-5 at 1/3 Cost Latent Space MiniMax news
2026-03-24 DeepSeek's Latest Job Postings Highlight Pivot to Agentic AI Bloomberg DeepSeek news
2026-03-23 SkillRouter Alibaba paper
2026-03-23 Felis ByteDance paper
2026-03-20 HuatuoGPT-3 CUHK model
2026-03-20 LongCat-Flash-Prover Meituan model 8564560B27B
2026-03-20 EnterpriseOps-Gym ServiceNow eval
2026-03-19 ★ Nemotron Cascade 2 NVIDIA model 3B1M1283/100
2026-03-19 dots.mocr Xiaohongshu model 3B
2026-03-18 Path-Constrained Mixture-of-Experts Apple paper
2026-03-18 Qianfan-OCR Baidu model 174.67K4B
2026-03-18 ★ MiniMax-M2.7 MiniMax model 2322/100
2026-03-18 MiMo-V2-Omni Xiaomi model 25
2026-03-18 ★ MiMo-V2-Pro Xiaomi model 1T42B1M29
2026-03-18 MiMo-V2-TTS Xiaomi model
2026-03-18 Chinese AI Developer Zhipu to Create New Unit for Product Development The Information Z.ai news
2026-03-17 PRISM: Demystifying Retention and Interaction in Mid-Training IBM paper
2026-03-17 ★ Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning SB Intuitions paper
2026-03-16 Mixture-of-Depths Attention (MoDA) ByteDance paper 267
2026-03-16 ★ Mistral Small 4 Mistral model 45.71K119B6.5B256K1139/100
2026-03-16 Attention Residuals 2 Moonshot AI paper library 3.3K
2026-03-16 CUBE: A Standard for Unifying Agent Benchmarks ServiceNow paper
2026-03-16 Bridging the Semantic Gap: Announcing the SEA-LION Embedding Suite AI Singapore AI Singapore news
2026-03-15 Scientific Judge 2 Baidu paper dataset 405
2026-03-15 Kernel Design Agents + KernelWiki MIT library
2026-03-13 OpenSWE / daVinci-Env SII dataset 187
2026-03-12 RoboBrain-Dex BAAI model 41
2026-03-12 IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL CMU, MBZUAI paper
2026-03-12 IndexShare (IndexCache): Cross-Layer Index Reuse for Sparse Attention Z.ai paper
2026-03-11 ★ Nemotron 3 Super NVIDIA model 745.73K120B12B1M1383/100
2026-03-10 ★ Exclusive Self Attention Apple paper
2026-03-10 Med-Asagi UTokyo model
2026-03-10 Ai2 CEO Ali Farhadi Steps Down; Microsoft Hires Key Researchers GeekWire Ai2 news
2026-03-09 Anthropic Sues Trump Admin Over Pentagon AI Blacklist CNBC Anthropic news
2026-03-09 OpenAI Acquires Promptfoo for AI Agent Security OpenAI OpenAI news
2026-03-08 CLI-Anything HKU library
2026-03-08 Scalable Training of MoE Models with Megatron Core NVIDIA paper
2026-03-06 ★ Sarvam-105B Sarvam model 12.57K105B10.3B128K939/100
2026-03-06 ★ Sarvam-30B Sarvam model 30B2.4B32K739/100
2026-03-05 ★ GPT-5.4 OpenAI model ~2.2T1M396/100
2026-03-05 GPT-5.4 Released with 1M Token Context OpenAI OpenAI news
2026-03-04 RIVER PJLab dataset 10
2026-03-01 ★ OLMo Hybrid Ai2 model 48.9K7B6T
2026-03-01 ★ LLM-jp-4 NII, University College London model 4.22K32B3.8B11.7T65.54K
2026-02-28 AnyTouch2 / ToucHD 2 BAAI dataset paper
2026-02-27 Solaris NYU model
2026-02-25 MaxClaw MiniMax library
2026-02-25 ZUNA (EEG Foundation Model) Zyphra model 380M
2026-02-25 Tencent-Backed AI Startup StepFun Is Said to Plan Hong Kong IPO Bloomberg StepFun news
2026-02-19 ★ Gemini 3.1 Pro Google model 1M306/100
2026-02-19 Gemini 3.1 Pro Released, Ties #1 on AA Intelligence Index Google DeepMind Google news
2026-02-18 Large-scale Online Deanonymization with LLMs ETH Zürich paper
2026-02-17 OLMix Ai2 paper
2026-02-17 Tiny Aya Cohere model 1.79K3.35B536/100
2026-02-17 ★ Grok-4.20 SpaceXAI model 2M26
2026-02-17 Mercury 2 Released: Diffusion LLM with AA Index 33 at 1000 tok/s Inception Labs Inception Labs news
2026-02-16 ★ Qwen3.5 5 Alibaba model 3.54K397B17B1M2139/100
2026-02-16 WebWorld Alibaba model 391.26K32B
2026-02-16 ★ Ling 2.5 Ant Group model 2151T1M
2026-02-16 ZoomBench Ant Group dataset 155
2026-02-15 Optimal Batch Size Scheduling via Functional Scaling Laws Meituan paper
2026-02-14 ★ Doubao-Seed-2.0 ByteDance model
2026-02-14 Doubao-Seed-2.0 Family Launched (Pro / Lite / Mini / Code) TechNode ByteDance news
2026-02-13 Cohere's $240M Year Sets Stage for IPO TechCrunch Cohere news
2026-02-12 ★ MiniMax-M2.5 MiniMax model 585589.2K229B2328/100
2026-02-12 GEBench StepFun dataset 54
2026-02-12 FireRed-Image-Edit Xiaohongshu model
2026-02-12 Xiaomi-Robotics-0 2 Xiaomi model paper 4.7B
2026-02-12 Anthropic Raises $30B Series G at $380B Valuation Anthropic Anthropic news
2026-02-11 Ming-Flash-Omni-2.0 Ant Group model 2.66K
2026-02-11 MiniCPM-SALA OpenBMB model 9.42K6.33K1M
2026-02-11 ★ Step-3.5-Flash 3 StepFun model paper dataset 2.08K325.86K196B11B256K17
2026-02-11 ★ GLM-5 3 Z.ai model paper 3.39K102.79K744B40B28.5T2850/100
2026-02-11 Slime: Asynchronous RL for Agentic Tasks Z.ai library 6.07K
2026-02-09 Protenix ByteDance model 1.94K
2026-02-09 InternAgent-1.5 PJLab paper
2026-02-08 ★ Data Darwinism / Darwin Corpora SII dataset 154
2026-02-07 Seedance 2.0 ByteDance model
2026-02-07 FireRed-OpenStoryline Xiaohongshu library
2026-02-06 Baichuan-M3 2 Baichuan paper model 2451.84K235B
2026-02-05 ★ Claude Opus 4.6 Anthropic model ~5.3T1M3211/100
2026-02-05 KernelGYM & Dr.Kernel HKUST model
2026-02-05 ★ Kling 3.0 2 Kuaishou model paper
2026-02-05 Claude Opus 4.6 Released with 1M Context Anthropic Anthropic news
2026-02-04 RationaleRM Alibaba dataset
2026-02-03 ★ MiniCPM-o 4.5 OpenBMB model 25.58K203.83K
2026-02-03 EchoJEPA Toronto & Vector Institute model
2026-02-02 WAXAL: African Language Speech Corpus Google dataset
2026-02-02 ★ Kimi K2.5 2 Moonshot AI model paper 2.02K1.64M32B2333/100
2026-02-02 daVinci-Agency SII model 9
2026-02-02 SpaceX Acquires xAI at $1.25T Combined Valuation Fortune SpaceXAI news
2026-02-01 nanobot HKU library
2026-01-30 Keel: Post-LayerNorm Is Back ByteDance paper
2026-01-30 GLM-OCR Z.ai model 900M
2026-01-29 SenseNova-MARS 3 SenseTime model paper dataset 1147432B (max)
2026-01-28 Qwen3-ASR Alibaba model 1.7B (max)
2026-01-28 ★ Trinity Large Arcee model 523398B13B17T512K
2026-01-28 Trinity Mini / Nano Arcee model 14.85K26B3B131K
2026-01-28 SDPO (Self-Distillation) ETH Zürich paper
2026-01-28 ACE-Step-1.5 StepFun model 10.97K
2026-01-27 ★ K2 Think V2 MBZUAI model 1.28K70B262.14K1189/100
2026-01-27 LongCat-Flash-Lite 2 Meituan model paper 2.08K1144/100
2026-01-27 Mistral AI Surges Revenue 20-Fold to Over $400 Million ARR MLQ Mistral news
2026-01-27 Tencent Bets Its AI Future on 28-Year-Old From OpenAI Caixin Tencent news
2026-01-26 DeepPlanning Alibaba dataset
2026-01-26 ★ daVinci-Dev SII model 28
2026-01-26 ★ Solar Pro 3 Upstage model 102B12B128K8
2026-01-23 ★ LongCat-Flash-Thinking-2601 2 Meituan model paper 2547.66K560B27B
2026-01-22 ★ ERNIE 5.0 2 Baidu model paper 22.4T14
2026-01-22 EvoCUA Meituan library 3234.12K
2026-01-21 CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning Alibaba eval
2026-01-21 The Flexibility Trap (JustGRPO) Tsinghua paper
2026-01-20 ★ Yuan 3.0 Ultra Inspur model 236351T68.8B
2026-01-20 Scale-RAE NYU model
2026-01-20 Step-3-VL-10B 2 StepFun model paper 406484.26K842/100
2026-01-16 Thinking Machines Suffers Wave of Defections as Co-Founders Zoph and Metz Return to OpenAI; Soumith Chintala Named CTO Fortune Thinking Machines news
2026-01-15 Francis Bach Awarded 2026 ERC Grant on the Reliability of AI Inria INRIA news
2026-01-15 Tao Qin Elected 2025 ACM Fellow ACM ZGCA news
2026-01-12 Engram: Conditional Memory via Scalable Lookup DeepSeek paper
2026-01-12 Alphabet Hits $4T Market Cap CNBC Google news
2026-01-11 Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models Baidu paper
2026-01-11 ★ Solar Open 100B Upstage model 3.87K102B12B19.7T131.07K1028/100
2026-01-09 Agnes-SeaLLM-8B Sapiens AI model 8B
2026-01-09 PaCoRe: Learning to Scale Test-Time Compute StepFun paper 334110
2026-01-09 Zhipu and MiniMax IPO ChinaTalk MiniMax news
2026-01-09 Zhipu and MiniMax IPO ChinaTalk Z.ai news
2026-01-08 Afri-MCQA MBZUAI eval
2026-01-06 One Sample to Rule Them All: Extreme Data Efficiency in Multidiscipline Reasoning with Reinforcement Learning Shanghai Jiao Tong University, Alibaba paper
2026-01-06 xAI Raises $20B Series E at $230B Valuation CNBC SpaceXAI news
2026-01-05 Yuan 3.0 Flash Inspur model 187840B3.7B
2026-01-05 ★ K-EXAONE LG model 236B23B1428/100
2026-01-05 ★ LFM2.5 Liquid AI model 8.3B1.5B38T128K733/100
2026-01-05 ★ HyperCLOVA X SEED Omni Naver model 6488B
2026-01-05 ★ Falcon-H1R TII model 95.76K7B256K844/100
2026-01-03 ★ HyperCLOVA X SEED Think Naver model 156.09K32B128K1131/100
2026-01-01 FlashInfer-python-paddle Baidu library
2026-01-01 Agentar-Z-100K Z.ai dataset
2025-12-31 FineWeb-Mask ByteDance dataset
2025-12-31 mHC: Manifold-Constrained Hyper-Connections DeepSeek paper
2025-12-31 OpenOneRec Kuaishou library 81233
2025-12-30 SeedFold ByteDance paper
2025-12-30 LongCat ZigZag Attention Meituan paper 842
2025-12-29 KV-Tracker Imperial paper
2025-12-27 RollArt: Disaggregated Multi-Task Agentic RL Training at Scale Alibaba paper
2025-12-27 ★ A.X K1 SK Telecom model 3033.17K519B33B10T131K
2025-12-23 ★ MiniMax-M2.1 MiniMax model 54410.19K229B2128/100
2025-12-23 VIBE & OctoCodingBench MiniMax dataset
2025-12-23 Step-DeepResearch StepFun library 561
2025-12-23 Zhipu AI's Rise from Tsinghua Lab Pandaily Z.ai news
2025-12-22 SekoTalk / Seko 2.0 SenseTime model 43
2025-12-22 ★ GLM-4.7 Z.ai model 2244/100
2025-12-20 TurboDiffusion Tsinghua library
2025-12-19 ★ Kanana-2 Kakao model 15030B3B32K
2025-12-19 Kakao Open-Sources Kanana-2 Model Optimized for Agentic AI Korea Times Kakao news
2025-12-18 ★ Seed1.8 ByteDance model 218
2025-12-18 EXAONE Path 2.5 LG paper
2025-12-18 Towards Scalable Pre-training of Visual Tokenizers MiniMax paper 490
2025-12-18 HY-Motion 1.0 Tencent paper
2025-12-18 Native Structured Latents for 3D Generation Tsinghua paper
2025-12-18 Seed1.8 Released as a Generalized Agentic Model ByteDance Seed ByteDance news
2025-12-17 Peter DeSantis to Lead Unified AGI Org; Rohit Prasad Departing CNBC Amazon news
2025-12-17 Tencent restructures AI operations, promotes high-profile recruit to chief AI scientist SCMP Tencent news
2025-12-16 ★ Molmo 2 Ai2 model 64118B (max)
2025-12-16 Motus Tsinghua model
2025-12-16 ★ MiMo-V2-Flash 2 Xiaomi model paper 1.33K70.62K309B15B27T2250/100
2025-12-16 MOPD (Multi-Teacher On-Policy Distillation) Xiaomi library 1.33K
2025-12-15 ★ Nemotron 3 Nano NVIDIA model 2.18M30B3.5B25T1M983/100
2025-12-15 SonicMoE Princeton library
2025-12-15 NVIDIA in Advanced Talks to Acquire AI21 Labs for $2-3B SiliconANGLE AI21 Labs news
2025-12-10 ★ LLaDA 2 2 Ant Group model paper 4309.38K
2025-12-09 ★ JAIS 2 MBZUAI model 2.5K70B2.6T8.19K
2025-12-08 LongCat-Image 3 Meituan model paper 69547.59K6B
2025-12-06 ★ K2-V2 (LLM360) MBZUAI model 18370B1089/100
2025-12-05 NEO (Native VLM Architecture) 2 SenseTime model paper 82519B (max)
2025-12-05 ★ Hunyuan 2.0 Tencent model 406B32B256K
2025-12-04 Nex-N1: Agentic Models via Large-Scale Environment Construction Nex-AGI paper
2025-12-04 Light-X NTU model
2025-12-03 Terminus-KIRA KRAFTON library
2025-12-02 ★ Amazon Nova 2 Amazon model 1M1411/100
2025-12-02 ★ Mistral Large 3 Mistral model 1.98K675B41B256K939/100
2025-12-02 Nova 2 Model Family and Nova Act GA at re:Invent 2025 TechCrunch Amazon news
2025-12-02 Anthropic Acquires Bun, Claude Code Hits $1B ARR Anthropic Anthropic news
2025-12-02 Thomson Reuters and Imperial Announce Frontier AI Research Lab, Targeting a Competitive Frontier LLM Imperial College London Imperial news
2025-12-01 Ministral 3 Mistral model 782.22K14B (max)539/100
2025-12-01 John Giannandrea to Retire; Amar Subramanya Named VP of AI Apple Apple news
2025-11-30 gelab-zero (STEP-GUI) StepFun library 2.19K2.21K
2025-11-28 ★ LFM2 (Liquid Foundation Models 2) Liquid AI model 24B2.3B628/100
2025-11-27 DeepSeek-Math-V2 2 DeepSeek model dataset 1.59K384
2025-11-27 MultiBanana UTokyo eval
2025-11-25 CABS + DataConcept-128M Tübingen dataset
2025-11-24 ★ Claude 4.5 Opus Anthropic model ~3.4T200K2911/100
2025-11-24 HunyuanOCR Tencent model 1.65K347.7K1B
2025-11-20 ★ OLMo 3 Ai2 model 10.31K32B5.9T65.54K789/100
2025-11-20 AICC: A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser PJLab dataset
2025-11-20 HunyuanVideo-1.5 Tencent model 4.47K2.34K
2025-11-20 MiMo-Embodied: X-Embodied Foundation Model Xiaomi paper 1.12K
2025-11-19 LPLB (Linear-Programming Load Balancer) DeepSeek library 505
2025-11-19 Step-Audio-R1 StepFun model 67322233B
2025-11-19 Yann LeCun Departs Meta to Found AMI Labs CNBC Meta news
2025-11-17 SenseNova-SI (Spatial Intelligence) 3 SenseTime model paper dataset 2718B (max)
2025-11-15 ★ Doubao Seed Code ByteDance model 256K1711/100
2025-11-15 Doubao Seed Code (Reasoning Coder) Hits AA Intelligence Index 34 Artificial Analysis ByteDance news
2025-11-14 Miloco (Xiaomi Local Copilot) Xiaomi library 2.61K
2025-11-13 M100 Chip Baidu announcement
2025-11-13 Murati's Thinking Machines in Funding Talks at $50 Billion Value Bloomberg Thinking Machines news
2025-11-12 ★ AlphaProof Google paper
2025-11-12 Interview: Ant Group's Open Model Ambitions Interconnects Ant Group news
2025-11-12 Chinese AI Prodigy Luo Fuli Joins Xiaomi as Industry Competition for Talent Heats Up South China Morning Post (via Yahoo Finance) Xiaomi news
2025-11-10 kosong Moonshot AI library 520
2025-11-10 RLVE: Scaling Up RL with Adaptive Verifiable Environments University of Washington, Ai2 paper
2025-11-09 ReProbe MBZUAI paper
2025-11-06 InfinityStar ByteDance model
2025-11-06 Cambrian-S NYU model
2025-11-06 Step-Audio-EditX StepFun model 92943.36K3B
2025-11-05 SoftBank and SB Intuitions launch Sarashina API for enterprise access to Japanese LLM SoftBank SB Intuitions news
2025-11-03 ★ HPLT 3 University of Edinburgh dataset
2025-11-03 ★ LongCat-Flash-Omni 2 Meituan model paper 49262560B27B
2025-11-01 popEVE Harvard model
2025-11-01 LightX2V SenseTime library 2.36K
2025-11-01 Inception Labs Raises $56M Seed from Menlo, Andrew Ng, Karpathy Inception Labs Inception Labs news
2025-10-31 GATE LG paper
2025-10-30 ★ Emu3.5 BAAI model 1.52K907
2025-10-30 Kimi Linear 2 Moonshot AI model paper 1.4K48B3B761/100
2025-10-30 ★ Nicheformer TU Munich model
2025-10-29 ★ Ouro ByteDance model 87.71K2.6B7.7T
2025-10-29 Toolathlon (The Tool Decathlon) HKUST eval 108
2025-10-29 ★ Gaperon INRIA model 24B4T
2025-10-28 ODesign BAAI model 311
2025-10-28 URSA (Uniform Discrete Diffusion) BAAI model
2025-10-28 Parallel Loop Transformer ByteDance paper
2025-10-28 OpenAI Completes For-Profit PBC Restructuring OpenAI OpenAI news
2025-10-27 ★ MiniMax-M2 MiniMax model 2.6K127.97K230B10B1928/100
2025-10-27 CoKE: Context as the Key to Biomolecular Understanding PJLab paper 18
2025-10-27 JanusCoder PJLab model 8047
2025-10-27 Hunyuan Mirror Tencent paper 1.14K2.99K1
2025-10-27 Artificial Hivemind University of Washington paper
2025-10-25 LongCat-Video 3 Meituan model paper 4.26K3.21K13.6B
2025-10-24 Huxley-Gödel Machine KAUST paper
2025-10-24 KAT-Coder 3 Kuaishou model paper 72B22
2025-10-23 Anthropic to Expand Google Cloud TPU Use to 1M+ TPUs Anthropic Anthropic news
2025-10-23 Aeronautics Spinout Monolith AI Acquired by CoreWeave Imperial College London Imperial news
2025-10-22 Seed3D 1.0 ByteDance model
2025-10-20 DeepSeek-OCR / OCR-2 DeepSeek model 23.27K2.35M
2025-10-17 LongCat-Audio-Codec Meituan paper 301
2025-10-16 MorphoBench ZGCA paper 13
2025-10-15 Claude Haiku 4.5 Anthropic model 200K1711/100
2025-10-15 ★ Granite 4.0 IBM model 32B9B656/100
2025-10-15 ★ ScaleRL: The Art of Scaling Reinforcement Learning Compute for LLMs Meta, UT Austin paper
2025-10-15 InteractiveOmni 2 SenseTime model paper 88B (max)
2025-10-15 Anthropic Releases Claude Haiku 4.5: Sonnet 4-Level Coding at One Third the Cost, $1/$5 per Million Tokens Anthropic Anthropic news
2025-10-15 Granite 4.0: Hybrid Mamba Architecture, First ISO 42001 Certified Open Models IBM IBM news
2025-10-14 Rex-Omni 2 IDEA Lab model paper 1.44K33.78K3B
2025-10-14 HAL — Holistic Agent Leaderboard Princeton eval
2025-10-14 Zhipu AI Breaks US Chip Reliance With First Major Model Trained on Huawei Stack SCMP Z.ai news
2025-10-13 RITE: Reinforcement Learning for Tool-Integrated Interleaved Thinking Meituan paper
2025-10-13 RAE — Representation Autoencoders NYU paper
2025-10-13 Rollout Routing Replay (R3) Xiaomi paper
2025-10-10 StreamingVLM MIT model
2025-10-10 DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning Sapiens AI paper
2025-10-09 ★ Ling 2.0 / Ling-1T 2 Ant Group model paper 3.35K1T50B944/100
2025-10-06 Paper2Video NUS library
2025-10-01 R-HORIZON-Websearch Meituan dataset 26
2025-10-01 Code2Video NUS library
2025-10-01 BroRL: Scaling Reinforcement Learning via Broadened Exploration NVIDIA paper
2025-10-01 GDPval OpenAI eval 11320
2025-10-01 ★ Apriel 15B Thinker ServiceNow model 131.07K1350/100
2025-10-01 Tinker Thinking Machines library
2025-10-01 IBM Research Names Jay Gambetta as Director; Dario Gil to DOE IBM IBM news
2025-09-30 Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation ByteDance paper
2025-09-30 ★ GLM-4.6 Z.ai model 355B1944/100
2025-09-29 Ring 4 Ant Group model paper 25838.2K1T63B262.14K1739/100
2025-09-29 ★ DeepSeek-V3.2 2 DeepSeek model paper 3.59M2685B37B21
2025-09-29 MGM-Omni HKUST model
2025-09-29 Scaling Behaviors of LLM Reinforcement Learning Post-Training PJLab paper
2025-09-29 AIRoA MoMa UTokyo dataset
2025-09-28 LLaVA-OneVision-1.5 NTU model
2025-09-28 HunyuanImage-3.0 2 Tencent model paper 3.12K2.61K13B
2025-09-26 Qwen3Guard Alibaba model 465
2025-09-25 Expanding Reasoning Potential (CoTP) Meituan paper
2025-09-24 LRM-Eval / ROME BAAI dataset 5
2025-09-23 ByteWrist ByteDance model
2025-09-23 ★ LongCat-Flash-Thinking 2 Meituan model paper 285103560B27B
2025-09-23 Symphony-MoE PCL paper
2025-09-22 BGE-Reasoner BAAI model 31959
2025-09-22 ScaleCUA PJLab model 1.11K58
2025-09-18 Seedream 4.0 ByteDance model 1
2025-09-17 ★ AToken Apple paper 140
2025-09-16 Shanghai launches innovation institute to bridge AI research and industry Shanghai Municipal Government SII news
2025-09-15 checkpoint-engine Moonshot AI library 963
2025-09-10 AgentGym-RL Fudan University paper
2025-09-10 Connectionism (research blog) Thinking Machines blog
2025-09-08 ★ mmBERT Johns Hopkins model 307M3T8.19K
2025-09-08 ★ PLaMo 2 PFN model 34.05K31B2T32K
2025-09-05 Klear 3 Kuaishou model paper 8258146B2.5B
2025-09-05 ★ MiniCPM4.1 2 OpenBMB model paper 9.42K49.98K8B
2025-09-02 Baichuan-M2 2 Baichuan paper model 212938132B
2025-09-02 Fantastic Pretraining Optimizers Stanford paper
2025-09-02 ★ Apertus Swiss AI model 161.34K270B15T65.54K589/100
2025-09-02 DynaGuard University of Maryland model
2025-09-02 Hugo Larochelle named Mila Scientific Director Mila Mila news
2025-09-01 VeOmni ByteDance library 2K
2025-09-01 Benchmarking Optimizers for LLM Pretraining EPFL paper
2025-09-01 Dream-Coder 7B & DreamOn HKU model
2025-09-01 ★ LongCat-Flash-Chat 2 Meituan model paper 1.34K81.37K1560B27B128K
2025-09-01 Hunyuan-MT Tencent model 71056.53K30B3B
2025-09-01 RLinf ZGCA library
2025-09-01 TwinBrainVLA ZGCA paper
2025-09-01 Darwin Monkey Zhejiang University announcement
2025-09-01 Mila unveils CAD $250M LaSalle Sovereign AI Research Hub Mila Mila news
2025-09-01 Mistral AI Raises EUR 2B at EUR 12B Valuation Mistral AI Mistral news
2025-08-28 HyperOS 3 Xiaomi announcement
2025-08-26 ★ MiniCPM-V 4.5 2 OpenBMB model paper 25.58K93.36K
2025-08-25 SEA-LION v4 AI Singapore model 32B (max)500B131.07K
2025-08-25 GEPO PCL paper
2025-08-25 InternVL 3.5 PJLab model 10.06K3241B (max)28B (max)
2025-08-25 SEA-LION v4: Our First Multimodal Release (Gemma-SEA-LION-v4-27B-IT, #5 of 55 on SEA-HELM) AI Singapore AI Singapore news
2025-08-23 HunyuanVideo-Foley Tencent paper
2025-08-21 Fin-PRM: Process Reward Model for Financial Reasoning Alibaba paper 502
2025-08-21 Waver ByteDance model 938
2025-08-21 ★ DeepSeek-V3.1 DeepSeek model 14
2025-08-21 Intern-S1 PJLab model 241B28B
2025-08-20 Seed-OSS-36B ByteDance model 88536.95K36B12T512K1244/100
2025-08-20 ShizhenGPT CUHK model
2025-08-20 Nemotron Nano V2 NVIDIA model 15.37K12B128K
2025-08-20 Seed-OSS-36B Released as Apache-2.0 Open-Weight Model VentureBeat ByteDance news
2025-08-19 NovoMolGen Mila model
2025-08-15 PXDesign ByteDance model 229
2025-08-15 Physical Autoregressive Model (PAR) PCL paper
2025-08-14 NextStep-1 2 StepFun model paper 6886214B
2025-08-14 Hunyuan-GameCraft 1.0 Tencent model 72376
2025-08-14 Cohere Raises $500M at $6.8B Valuation Cohere Cohere news
2025-08-14 Cohere Hires Long-Time Meta Research Head Joelle Pineau as Chief AI Officer TechCrunch Cohere news
2025-08-12 OpenCUA HKU model
2025-08-12 Mistral Medium 3.1 Mistral model 128K911/100
2025-08-12 InternBootcamp PJLab library 349
2025-08-11 GLM-4.5V Z.ai model 2.33K167.01K106B12B64K853/100
2025-08-07 LLMEval-Fair Fudan University eval 200K
2025-08-07 CANN Huawei library
2025-08-07 ★ GPT-5 OpenAI model ~4.1T400K236/100
2025-08-07 TMA-Adaptive FP8 Grouped GEMM PJLab paper 25
2025-08-06 ACAVCaps Xiaomi dataset 424
2025-08-05 OmniScale ByteDance paper 2K
2025-08-05 Seed Diffusion ByteDance model
2025-08-05 ★ gpt-oss OpenAI model 117B5.1B131.07K1239/100
2025-08-05 dots.vlm1 Xiaohongshu model
2025-08-05 OpenAI Releases gpt-oss-120b and gpt-oss-20b, Its First Open-Weight Language Models Since GPT-2 (Apache 2.0) Hugging Face OpenAI news
2025-08-01 Qwen-Image 3 Alibaba model 7.98K173.33K20B
2025-08-01 MegaDFT ZGCA paper
2025-08-01 Ai2 and UW Awarded $152M from NSF and NVIDIA for Open Scientific AI GeekWire Ai2 news
2025-07-31 Seed-Prover ByteDance model 433
2025-07-30 dots.ocr Xiaohongshu model 3B
2025-07-29 Libra-Bench & PIE_bench Meituan dataset
2025-07-28 SmallThinker Shanghai Jiao Tong University model 21B3B
2025-07-28 MixGRPO Tencent paper 1.14K
2025-07-28 ★ GLM-4.5 2 Z.ai model paper 2355B1356/100
2025-07-27 SenseNova V6.5 SenseTime model
2025-07-27 StepFun-Prover-Preview StepFun model 356032B
2025-07-27 HunyuanWorld 3 Tencent model 2.85K615
2025-07-26 ARPO (Agentic Reinforced Policy Optimization) Renmin (RUC) paper
2025-07-25 ★ Step-3 2 StepFun model paper 453144.09K321B38B
2025-07-24 A.X 3.1 SK Telecom model 1330534B
2025-07-24 SoftBank Corp. to Build the World's Largest AI Computing Infrastructure Using NVIDIA DGX SuperPOD with NVIDIA Blackwell GPUs SB Intuitions SB Intuitions news
2025-07-23 Towards Greater Leverage: Scaling Laws for Efficient MoE Ant Group paper
2025-07-23 ASI-Arch SII paper 726
2025-07-22 Qwen-Code Alibaba library 25.07K
2025-07-22 Qwen3-Coder 2 Alibaba model 16.61K1M480B (max)35B (max)1044/100
2025-07-22 Seed-X Series ByteDance model 1721307B
2025-07-22 Reka Raises $110M Series B at $1B Valuation Reka Reka news
2025-07-17 Agentar-DeepFinance-100K Ant Group dataset 35
2025-07-17 Apple Intelligence Foundation Models Tech Report 2025 Apple Apple news
2025-07-15 Ettin Johns Hopkins model 1B2T
2025-07-15 Mira Murati's Thinking Machines Lab Is Worth $12B in Seed Round TechCrunch Thinking Machines news
2025-07-14 HKGAI-V1 HKUST model
2025-07-14 Mixture-of-Recursions KAIST paper
2025-07-14 ★ EXAONE 4.0 LG model 33.89K32B828/100
2025-07-12 Scaling Laws for Optimal Data Mixtures Apple paper
2025-07-11 ★ Kimi K2 4 Moonshot AI model paper 10.84K2.71M51T32B256K1344/100
2025-07-10 FlexOlmo Ai2 model 1501.4K33B
2025-07-10 H-Net CMU model
2025-07-10 KAT (Kwai-AutoThink) 2 Kuaishou paper model 5572B
2025-07-09 EXAONE Path 2.0 LG paper
2025-07-09 CADmium Mila model
2025-07-09 ★ Grok-4 SpaceXAI model ~3.2T256K226/100
2025-07-09 Grok-4 Released with Native Tool Use and Reasoning xAI SpaceXAI news
2025-07-08 ★ NeoBabel University of Amsterdam model
2025-07-07 POLAR PJLab paper 167
2025-07-05 How to Train Your LLM Web Agent ServiceNow paper
2025-07-03 ★ IFBench Ai2 eval 140130058
2025-07-03 ★ A.X 4.0 SK Telecom model 158653
2025-07-01 CodePRM: Execution Feedback-enhanced Process Reward Model for Code Generation Huawei paper
2025-07-01 Voxtral Mistral model 337.47K
2025-07-01 mini-swe-agent Princeton library
2025-07-01 ★ Solar Pro 2 Upstage model 31B64K711/100
2025-06-30 ★ openPangu Huawei announcement
2025-06-30 Meta Superintelligence Labs Created; Wang Named Chief AI Officer CNBC Meta news
2025-06-27 ★ HyperCLOVA X THINK Naver model 128K
2025-06-27 Hunyuan-A13B 2 Tencent model paper 81646.54K80B13B256K
2025-06-26 Kwai Keye-VL 3 Kuaishou model paper 785192.8K31B3B262.14K
2025-06-25 OctoThinker SII model 1888B (max)
2025-06-24 Video-XL-2 BAAI model 54
2025-06-24 Radial Attention MIT paper
2025-06-23 The Open Proof Corpus ETH Zürich dataset
2025-06-18 Show-o2 NUS model 7B
2025-06-18 Thunder-DeID SNU dataset
2025-06-18 ★ Thunder-LLM SNU model
2025-06-18 Thunder-Tok SNU paper
2025-06-17 ★ Mercury (Diffusion LLM) Inception Labs model 1128K14
2025-06-17 PFMBench: Protein Foundation Model Benchmark Westlake University eval 38
2025-06-16 SciSage / SurveyScope BAAI library
2025-06-16 ★ MiniMax-M1 2 MiniMax model paper 3.15K855456B45.9B1M12
2025-06-15 MOSS-TTSD Fudan University model
2025-06-15 ★ AI-Driven Agentic Design Platform for Tumor Immunotherapy Drugs ZGCA announcement
2025-06-15 PR[AI]RIE-PSAI AI Cluster Wins €75 Million in France 2030 Funding Université PSL INRIA news
2025-06-15 ZGCA & ZGCI Unveil AI-Driven Tumor Immunotherapy Drug Design Platform Zhongguancun Academy ZGCA news
2025-06-14 OpenUnlearning CMU library
2025-06-13 Scientists' First Exam PJLab eval 83066
2025-06-12 Seed-1.6 (AdaCoT) ByteDance model 256K
2025-06-12 ★ Magistral Mistral model 38.49K224B (max)950/100
2025-06-12 Predictable Scale Part II: Farseer StepFun paper
2025-06-12 Seed-1.6 Introduces Adaptive Chain-of-Thought (AdaCoT) ByteDance Seed ByteDance news
2025-06-11 FlagEvalMM BAAI library 106
2025-06-11 Attention-Based Map Encoding (AME) ETH Zürich paper
2025-06-10 Seedance 1.0 ByteDance model
2025-06-09 RedNote joins AI race with its own open-source model that it says bests Alibaba, DeepSeek SCMP Xiaohongshu news
2025-06-07 ★ The Illusion of Thinking Apple paper
2025-06-07 Polaris HKU model
2025-06-06 RoboBrain 2.0 2 BAAI model 1.09K573
2025-06-06 ★ MiniCPM4 2 OpenBMB model paper 9.42K11.73K8B
2025-06-06 Ultra-FineWeb OpenBMB dataset 30
2025-06-06 ★ dots.llm1 2 Xiaohongshu model paper 142B14B11.2T32.77K
2025-06-05 RoboRefer / RefSpatial BAAI model 263
2025-06-05 Boltzmann Sampler Line (PTSD / BNEM / DiKL) University of Cambridge paper
2025-06-05 Loss Deceleration & Zero-Sum Learning Mila paper
2025-06-05 Log-Linear Attention MIT paper
2025-06-05 The Common Pile 2 Toronto & Vector Institute dataset model 7B
2025-06-04 EuroLLM University of Edinburgh model 22B (max)
2025-06-04 MiMo-VL 2 Xiaomi model paper 6433.93K7B
2025-06-04 Chinese social media app Xiaohongshu's $26 billion valuation bolsters GSR fund Bloomberg Xiaohongshu news
2025-06-03 CyberGym UC Berkeley eval
2025-06-03 KRONOS Harvard model
2025-06-03 UniWorld Peking University model
2025-06-02 OWSM / OWLS CMU model 18B
2025-06-01 HumanSense Benchmark Ant Group dataset
2025-06-01 BrowseComp & WideSearch Moonshot AI dataset
2025-06-01 kimi-agent-sdk Moonshot AI library 487
2025-06-01 kimi-cli Moonshot AI library 8.94K
2025-06-01 Kimi-Dev 2 Moonshot AI model paper 1.23K2.67K72B
2025-06-01 Kimi-Researcher Moonshot AI model 81
2025-06-01 walle Moonshot AI library 21
2025-06-01 AgentCPM Series 3 OpenBMB paper 8001.67K
2025-06-01 A.X Encoder SK Telecom model 2.68K
2025-06-01 CF-Div2-Stepfun StepFun dataset
2025-06-01 SteptronOss StepFun library 575
2025-06-01 Cryo-IEF Westlake University model
2025-06-01 MiMo-Audio 2 Xiaomi model paper 7B
2025-06-01 BrowseComp Z.ai dataset
2025-05-30 AReaL Ant Group library 5.29K
2025-05-30 ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries NVIDIA paper
2025-05-29 MathArena ETH Zürich eval
2025-05-29 KVzip SNU paper
2025-05-29 BioReason Toronto & Vector Institute model
2025-05-29 TrajViT University of Washington paper
2025-05-28 Ming-Omni Ant Group model 65642
2025-05-28 ★ DeepSeek-R1-0528 DeepSeek model 671B37B
2025-05-28 Pangu Embedded Huawei paper 118
2025-05-28 ★ Skywork Open Reasoner 1 Skywork model 74332B
2025-05-27 Pangu Pro MoE Huawei paper 1331
2025-05-27 GTA & GLA — Hardware-Efficient Attention Princeton paper
2025-05-27 HunyuanVideo-Avatar Tencent paper
2025-05-26 FLAME-MoE CMU model
2025-05-26 SynLogic 2 MiniMax paper dataset 203121
2025-05-26 VLM-3R UT Austin paper
2025-05-26 VoiceStar UT Austin model
2025-05-25 ★ FP4 All the Way Technion paper
2025-05-24 AI-Researcher HKU library
2025-05-23 Why Diffusion Models Don't Memorize INRIA paper
2025-05-23 One RL to See Them All: Visual Triple Unified RL MiniMax paper 333
2025-05-23 FairyR1-32B Peking University model
2025-05-22 ★ Claude 4 Anthropic model ~1.4T1M2111/100
2025-05-22 The Polar Express NYU paper
2025-05-22 XRing O1 Xiaomi announcement
2025-05-21 Reverse Engineering Human Preferences with RL Imperial paper
2025-05-21 Devstral 2 Mistral model 123B256K928/100
2025-05-21 ★ Falcon-H1 TII model 11811.02K34B256K
2025-05-20 BAGEL ByteDance model 6K82014B
2025-05-20 s3 UIUC paper
2025-05-19 On the Conversational Persuasiveness of GPT-4 EPFL paper
2025-05-19 ★ Marin Stanford model 32B
2025-05-19 Quetzal Toronto & Vector Institute model
2025-05-18 mCLM UIUC paper
2025-05-17 Video-SafetyBench BAAI eval 122.26K
2025-05-17 Model Merging in Pre-training of LLMs ByteDance paper
2025-05-15 BGE-Code-v1 BAAI model 11.8K4.84K
2025-05-15 MMLongBench University of Edinburgh eval 13.33K
2025-05-15 Superposition Yields Robust Neural Scaling MIT paper
2025-05-15 Apriel-Nemotron-15B Reasoning Model with NVIDIA ServiceNow ServiceNow news
2025-05-14 ★ AlphaEvolve Google paper 6
2025-05-12 Seed1.5-VL ByteDance model 1.58K200B20B
2025-05-12 MiniMax-Speech: Intrinsic Zero-Shot TTS MiniMax paper
2025-05-12 Step1X-3D: High-Fidelity Textured 3D Assets StepFun model 870
2025-05-11 GuidedQuant SNU paper
2025-05-10 Gated Attention for Large Language Models Alibaba paper
2025-05-08 Seed-Coder-8B ByteDance model 7558B65.54K
2025-05-07 DeerFlow ByteDance library 70.89K
2025-05-07 ★ Pangu Ultra MoE 2 Huawei model paper 16718B39B
2025-05-07 General-Level & General-Bench NUS eval 700
2025-05-07 HunyuanCustom Tencent paper 1.22K1
2025-05-06 CCI 4.0 BAAI dataset
2025-05-06 OpenSeek BAAI model 2614
2025-05-06 RoboOS 2 BAAI library 575
2025-05-06 SkyRL UC Berkeley library
2025-05-06 VideoMimic UC Berkeley paper
2025-05-06 WebGen-Bench CUHK eval 647101
2025-05-06 Absolute Zero Reasoner Tsinghua paper
2025-05-06 OpenHelix Westlake University model
2025-05-05 El Agente Q Toronto & Vector Institute paper
2025-05-02 ★ MiMo (Reasoning) 2 Xiaomi model paper 91.71K17B
2025-05-01 AWorld Ant Group library 1.2K
2025-05-01 ★ Kanana 1.5 Kakao model 8915.7B3B32K
2025-05-01 M2PDE / RealPDEBench Westlake University paper
2025-04-30 ★ Amazon Nova Premier Amazon model ~470B1M911/100
2025-04-30 DeepSeek-Prover-V2 DeepSeek model 1.27K633671B
2025-04-30 SWE-smith Princeton dataset
2025-04-30 WebThinker Renmin (RUC) paper
2025-04-30 Nova Premier Launched as Amazon's Most Capable AI Model TechCrunch Amazon news
2025-04-30 Phi-4 Reasoning Models Released with Chain-of-Thought Microsoft Microsoft news
2025-04-29 ★ Qwen3 9 Alibaba model paper 27.3K741T22B (max)9
2025-04-29 Behavior-SD SNU dataset
2025-04-29 First LlamaCon Developer Conference Meta AI Meta news
2025-04-27 Explanatory Summarization with Discourse-Driven Planning University of Edinburgh paper
2025-04-25 PolyMath Alibaba dataset 44
2025-04-25 Kimi-Audio 2 Moonshot AI model paper 4.65K80.22K7B
2025-04-24 gRNAde University of Cambridge model
2025-04-24 The Sparse Frontier University of Edinburgh paper
2025-04-24 Paper2Code KAIST library
2025-04-24 Step1X-Edit StepFun model 2.22K110
2025-04-22 LiveCC NUS model
2025-04-22 TTRL: Test-Time Reinforcement Learning PJLab paper
2025-04-21 LUFFY: Learning to Reason under Off-Policy Guidance Westlake University paper
2025-04-19 SRPO: Staged History-Resampling Policy Optimization Kuaishou paper
2025-04-18 Does RL Really Incentivize Reasoning Beyond the Base Model? Tsinghua paper
2025-04-17 Nemotron-CLIMB: Clustering-based Iterative Data Mixture Bootstrapping 3 NVIDIA paper dataset 1
2025-04-16 ★ o3 OpenAI model ~3T200K206/100
2025-04-15 DataDecide Ai2 paper
2025-04-15 ReTool: Reinforcement Learning for Strategic Tool Use in LLMs ByteDance paper 2
2025-04-15 ★ Kling 2.0 1 Kuaishou model
2025-04-15 Kimina-Prover 2 Moonshot AI model paper 370717
2025-04-15 miniF2F-test (Rectified) Moonshot AI dataset 370
2025-04-15 ★ Apriel ServiceNow model 32415B (max)4.5T
2025-04-15 Step-R1-V-Mini StepFun model
2025-04-15 ZR1-1.5B Zyphra model 1.5K1.5B
2025-04-15 Apriel-5B: ServiceNow's First Open SLM ServiceNow ServiceNow news
2025-04-14 SocioVerse Fudan University library
2025-04-14 InternVL3 PJLab model 10.06K578B (max)
2025-04-12 SenseNova V6 SenseTime model 600B
2025-04-11 CellFlow TU Munich paper
2025-04-10 ★ Scaling Laws for Native Multimodal Models Apple paper
2025-04-10 ★ Seed1.5-Thinking: Advancing Superb Reasoning Models with RL ByteDance paper 1
2025-04-10 ★ Pangu Ultra 2 Huawei model paper 77135B
2025-04-10 Kimi-VL 2 Moonshot AI model paper 1.2K123.09K116B2.8B
2025-04-09 SoK: Membership Inference Attacks on LLMs Imperial paper
2025-04-09 DisCIPL (Self-Steering LMs) MIT paper
2025-04-09 A Sober Look at Progress in LM Reasoning Tübingen paper
2025-04-08 Amazon Nova Sonic Amazon model
2025-04-08 Dream 7B HKU, Huawei model 1.25K7B
2025-04-08 ★ Skywork R1V Series Skywork model 4138B
2025-04-08 Nova Sonic Speech-to-Speech Model Launched on Bedrock AWS Amazon news
2025-04-07 BaichuanMed-OCR Baichuan model 16972B (max)
2025-04-07 EvoTune EPFL paper
2025-04-06 Metamon UT Austin paper
2025-04-05 ★ Llama 4 Meta model 400B17B1M1028/100
2025-04-05 Llama 4 Scout and Maverick Released (First MoE, Multimodal) Meta AI Meta news
2025-04-04 ★ Nemotron-H NVIDIA model 58.35K56B20T
2025-04-04 DeepResearcher Shanghai Jiao Tong University model 7B
2025-04-04 MedSAM2 Toronto & Vector Institute model
2025-04-03 DeepSeek-GRM: Inference-Time Scaling for Generalist Reward Modeling DeepSeek paper 1
2025-04-02 ATOMICA Harvard model
2025-04-01 MiniMax Speech Series MiniMax model
2025-03-31 Amazon Nova Act Amazon model 909
2025-03-31 Open-Reasoner-Zero: Scaling Up RL on the Base Model StepFun, Tsinghua paper
2025-03-31 Amazon Unveils Nova Act, an AI Agent That Controls a Web Browser TechCrunch Amazon news
2025-03-30 ★ ToRL: Scaling Tool-Integrated RL SII paper 3491
2025-03-28 Doubao-Deep-Thinking ByteDance model
2025-03-27 OpenComplex 2 BAAI model 269
2025-03-27 OlymMATH Renmin (RUC) eval
2025-03-26 Qwen2.5-Omni-7B Alibaba model 4.02K777.76K7B
2025-03-25 ★ Gemini 2.5 Pro Google model ~1.2T1M166/100
2025-03-24 SimpleRL-Zoo: Investigating and Taming Zero RL for Open Base Models ByteDance, Meituan paper
2025-03-24 CaMeL ETH Zürich paper
2025-03-24 ★ SimpleRL / SimpleRL-Zoo HKUST model
2025-03-22 Safe RLHF-V Peking University paper
2025-03-21 Hunyuan-T1 Tencent model
2025-03-21 CVE-Bench UIUC eval
2025-03-21 SEA-HELM: Assessing LLM Performance for Southeast Asia AI Singapore AI Singapore news
2025-03-20 ★ Aardvark Weather University of Cambridge model
2025-03-18 Sable BAAI model 2
2025-03-18 Isaac GR00T NVIDIA model
2025-03-18 ★ Llama-Nemotron (Nano/Super/Ultra) NVIDIA model 123.53K1128K853/100
2025-03-18 1000-Layer Networks for Self-Supervised RL Princeton paper
2025-03-18 HaploVL Tencent model 65
2025-03-17 DAPO 2 ByteDance paper dataset
2025-03-17 ★ EXAONE Deep LG model 1.4K32B
2025-03-17 SuperBPE University of Washington paper
2025-03-16 ★ ERNIE 4.5 Baidu model 7.72K424B (max)47B (max)856/100
2025-03-16 ERNIE X1 Baidu model
2025-03-14 TxAgent & ToolUniverse Harvard library
2025-03-14 VGGT University of Oxford, Meta model
2025-03-13 MMLU-ProX UTokyo eval
2025-03-12 ★ Gemma 3 Google model 1.42M4027B128K550/100
2025-03-12 Search-R1 UIUC paper
2025-03-11 TinyDeepSeek CUHK model 3.3B
2025-03-10 Seedream 2.0 ByteDance paper
2025-03-10 YOLOE Tsinghua model
2025-03-10 Reka Flash 3 Released (Open-Weight, 21B) Reka Reka news
2025-03-07 ★ Ling 2 Ant Group paper model 25838.63K116B
2025-03-07 R1-Searcher Renmin (RUC) paper
2025-03-06 QwQ-32B Alibaba model 52.45K32B9
2025-03-06 BGE-VL 2 BAAI model dataset 11.8K4.55K
2025-03-06 L1 (Length Controlled Policy Optimization) CMU paper
2025-03-06 Predictable Scale Part I: Step Law StepFun paper 1
2025-03-05 EgoLife & EgoGPT NTU model
2025-03-05 DINGO-BNS Tübingen paper
2025-03-04 CogView-4 Z.ai model 1.1K
2025-03-03 ★ Aya Vision Cohere model 160.27K32B
2025-03-01 ★ Command A Cohere model 2.28K111B256K733/100
2025-03-01 ★ Sarashina2.2 3 SB Intuitions model 12.98K4B
2025-02-28 3FS (Fire-Flyer File System) DeepSeek library 9.96K
2025-02-28 Smallpond DeepSeek library 4.96K
2025-02-28 Image-01 MiniMax model
2025-02-27 RoboBrain BAAI model 553132
2025-02-27 UniTok ByteDance paper 526
2025-02-27 DualPipe DeepSeek library 2.96K
2025-02-27 EPLB (Expert Parallelism Load Balancer) DeepSeek library 1.39K
2025-02-27 Hunyuan Turbo S 2 Tencent model paper 560B56B
2025-02-27 moscot TU Munich library
2025-02-26 DeepGEMM DeepSeek library 7.36K
2025-02-26 BIG-Bench Extra Hard (BBEH) Google eval 4.52K23
2025-02-26 ★ Kanana Kakao model 280132.5B3T
2025-02-26 NeoBERT Mila model 250M
2025-02-26 Granite 3.2: Multimodal Vision and Chain-of-Thought Reasoning IBM IBM news
2025-02-25 DeepEP DeepSeek library 9.71K
2025-02-25 Rank1 Johns Hopkins model
2025-02-24 ★ Claude Code Anthropic library 131.53K
2025-02-24 Baichuan-Audio 2 Baichuan paper model 22210310B
2025-02-24 FlashMLA DeepSeek library 12.7K
2025-02-24 Reasoning with Latent Thoughts: On the Power of Looped Transformers Google paper 1
2025-02-24 Muon Optimizer 2 Moonshot AI paper library 1
2025-02-24 Topic Over Source: The Key to Effective Data Mixing for LLM Pre-training PJLab paper 1
2025-02-24 AISafetyLab Tsinghua library
2025-02-24 Erwin University of Amsterdam paper
2025-02-22 Moonlight-3B/16B Moonshot AI model 1.49K32.57K116B (max)
2025-02-21 LightThinker Zhejiang University paper
2025-02-20 SEA-HELM AI Singapore eval
2025-02-20 ★ SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines ByteDance eval 188426.53K285
2025-02-20 Measuring CoT Faithfulness by Unlearning Reasoning Steps Technion paper
2025-02-19 Qwen2.5-VL Alibaba model 52
2025-02-19 FlexTok Apple paper 319
2025-02-19 MCQA critique: A) Forced B) Flawed C) Fixable D) All of the Above University of Maryland paper
2025-02-18 MoBA: Mixture of Block Attention for Long-Context LLMs Moonshot AI paper 2.13K2
2025-02-18 Hunyuan-Large-Vision Tencent model 389B52B
2025-02-18 CuatroLLM / TransWebLLM University College London model 1.3B
2025-02-17 Mistral Saba Mistral model 24B32K6
2025-02-17 OpenDWM / MaskGWM 2 SenseTime library paper 398
2025-02-17 ★ Grok-3 SpaceXAI model ~2.1T1M12
2025-02-17 Step-Audio / Step-Audio2 StepFun model 27
2025-02-17 Grok-3 Launched, Trained on 200K GPU Colossus Cluster xAI SpaceXAI news
2025-02-16 AdaGC: Improving Training Stability for Large Language Model Pretraining Baidu paper
2025-02-16 NSA: Native Sparse Attention DeepSeek paper 2
2025-02-15 1bit-Merging 1 Huawei paper
2025-02-14 WebOrganizer Ai2 paper
2025-02-14 ★ LLaDA 2 Ant Group model 3.82K2.27K5
2025-02-14 FineWeb2-HQ / FineWeb-HQ EPFL dataset
2025-02-14 KernelBench Stanford eval
2025-02-14 Step-Video-T2V 2 StepFun model paper 3.19K300B
2025-02-13 LOBS5 / LOB-Bench University of Oxford model
2025-02-12 WorldGUI & GUI-Thinker NUS eval
2025-02-12 AgentSociety Tsinghua library
2025-02-12 Wu Yonghui Joins ByteDance as Head of Seed Basic Research SCMP ByteDance news
2025-02-11 Scion EPFL paper
2025-02-11 CodeI/O HKUST dataset
2025-02-11 ★ Nature Language Model (NatureLM) Microsoft model 46146.7B13B
2025-02-11 TransMLA Peking University paper
2025-02-11 Goedel-Prover Princeton model
2025-02-10 Agentica RL Line (DeepScaleR / DeepCoder / DeepSWE) UC Berkeley model
2025-02-10 QServe / LServe (OmniServe) MIT library
2025-02-09 AutoAgent HKU library
2025-02-08 TabICL / CARTE INRIA model
2025-02-07 ★ Huginn-3.5B Tübingen, University of Maryland model 3.5B800B
2025-02-07 ★ Gemstones University of Maryland model 2B
2025-02-06 Great Models Think Alike Tübingen paper
2025-02-05 ★ Scaling Laws for Upcycling Mixture-of-Experts Language Models SB Intuitions paper 9
2025-02-05 ★ LIMO SII paper 543
2025-02-04 OpenAI and Kakao to Jointly Develop AI Products for South Korea CNBC Kakao news
2025-02-03 TwinMarket CUHK library
2025-02-02 Sundial / Timer-XL Tsinghua model
2025-02-01 ModernBERT-Ja SB Intuitions model 44.81K310M
2025-01-31 LR Schedules and Convex Optimization INRIA paper
2025-01-31 s1: Simple Test-Time Scaling Stanford paper
2025-01-30 MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding PJLab paper
2025-01-30 MedXpertQA PJLab eval 4.46K
2025-01-29 International AI Safety Report 2 Mila paper
2025-01-28 Over-Tokenized Transformer ByteDance paper
2025-01-28 THREADS Harvard model
2025-01-27 Training Dynamics of In-Context Learning in Linear Attention University College London paper
2025-01-26 Baichuan-Omni-1.5 2 Baichuan paper model 1911.84K7B
2025-01-24 Baichuan-M1 3 Baichuan paper model 821514.5B
2025-01-24 FireRedASR Xiaohongshu model 8.3B (max)
2025-01-23 Video-MMMU NTU eval 900
2025-01-23 UltraRAG OpenBMB library 5.58K
2025-01-22 ★ Doubao-1.5-Pro ByteDance model 200B20B256K
2025-01-22 UI-TARS ByteDance library 10.9K621.09K5
2025-01-22 ★ DeepSeek-R1 DeepSeek model 5.35M671B37B1350/100
2025-01-22 Revisit Self-Debugging with Self-Generated Tests for Code Generation Meituan paper 1
2025-01-21 TokenVerse Technion paper
2025-01-21 Hunyuan3D 2.0 3 Tencent model paper 13.93K75.95K6
2025-01-21 Stargate Project: $500B AI Infrastructure Initiative OpenAI OpenAI news
2025-01-20 ★ Kimi k1.5 2 Moonshot AI model paper 3.47K11
2025-01-17 ComplexFuncBench Z.ai dataset 180
2025-01-15 ★ SpeechGPT-2.0-preview Fudan University model 7B
2025-01-15 UNI2-h & CONCH v1.5 Harvard model
2025-01-15 ★ InternLM3 PJLab model 7.22K
2025-01-14 ★ MiniMax-01 3 MiniMax model paper 3.43K172.55K1456B4M
2025-01-14 ★ MiniCPM-o 2.6 OpenBMB model 25.58K386.43K8B
2025-01-14 InstructCell Zhejiang University model
2025-01-10 ★ Sky-T1 / SkyThought UC Berkeley model
2025-01-10 GThinker PCL model
2025-01-09 Search-o1 Renmin (RUC) paper
2025-01-09 WanJuan 3.0 (WanJuan-SiLu) PJLab dataset
2025-01-09 Asagi UTokyo model
2025-01-07 ★ Cosmos NVIDIA model 9
2025-01-06 Evolla Westlake University model 10B (max)
2025-01-03 AgentRefine Meituan paper
2025-01-02 SEDD (GPU deduplication) SNU paper
2025-01-02 FlashInfer University of Washington library
2025-01-01 Document Parse Upstage library
2024-12-31 ★ OLMo 2 Ai2 model 3.72K232B6T4.1K689/100
2024-12-30 SWE-Gym UC Berkeley dataset
2024-12-26 ★ DeepSeek-V3 3 DeepSeek model paper 227671B37B8
2024-12-25 QVQ Alibaba model
2024-12-25 ★ HuatuoGPT-o1 CUHK model
2024-12-24 ★ LLM-jp-3 (172B) NII model 20172B2.1T4.1K
2024-12-23 Baichuan4-Finance 2 Baichuan paper model 1
2024-12-20 Align-Anything Peking University library
2024-12-19 SLAM-LLM Shanghai Jiao Tong University library
2024-12-19 DreMa University of Amsterdam paper
2024-12-18 NOVA (Non-quantized Video Autoregressive) BAAI model 651
2024-12-18 TheAgentCompany CMU eval
2024-12-18 VSI-Bench NYU eval
2024-12-16 ★ MASt3R-SLAM Imperial library
2024-12-13 DeepSeek-VL2 DeepSeek model 5.3K2.68K22
2024-12-13 Liquid AI Raises $250M Series A Led by AMD Liquid AI Liquid AI news
2024-12-13 Profile: Shanghai AI Lab: Driving both AI safety and development MERICS PJLab news
2024-12-12 ★ Phi-4 Microsoft model 801.85K2414B650/100
2024-12-12 Phi-4 Released: 14B SLM Specializing in Complex Reasoning Microsoft Research Microsoft news
2024-12-11 SEA-LION v3 AI Singapore model 70B
2024-12-11 FlowEdit Technion paper
2024-12-10 ProCyon Harvard model
2024-12-09 ProcessBench Alibaba dataset 1
2024-12-07 SNU Korean Eval Suite SNU eval
2024-12-06 ★ Aya Expanse Cohere model 48.1K332B128K
2024-12-06 ★ EXAONE 3.5 LG model 10.25K232B
2024-12-06 Densing Law of LLMs OpenBMB paper 2
2024-12-06 InternVL 2.5 PJLab model 10.06K3491378B (max)
2024-12-06 APOLLO UT Austin library
2024-12-05 Language Model Ladders Ai2 paper
2024-12-05 Infinity & InfinityStar ByteDance model 1.57K1
2024-12-05 Liquid: Scalable Multi-modal Generation ByteDance model 6431
2024-12-05 Divot Tencent model 87
2024-12-05 Moto Tencent paper 177
2024-12-04 AceGPT & Native Alignment CUHK model
2024-12-04 GenCast Google paper
2024-12-04 RedStone Microsoft dataset 161
2024-12-03 Amazon Nova Amazon model 2~90B300K711/100
2024-12-03 AWS Trainium2 (Trn2 / Trn2 UltraServer) Amazon announcement
2024-12-03 HunyuanVideo Tencent model 12.19K613B
2024-12-03 SEED-Voken Tencent paper 1.01K
2024-12-03 GLM-4-Voice: End-to-End Spoken Chatbot Z.ai model 2
2024-12-03 AWS Trainium2 Chips Generally Available; Trainium3 Previewed TechCrunch Amazon news
2024-12-01 ★ flash-linear-attention MIT library
2024-12-01 ★ Falcon 3 TII model 10B
2024-12-01 ScanNet++ (v2) TU Munich dataset
2024-12-01 Rainbow Teaming University College London, Meta paper
2024-11-29 ★ TITAN Harvard model
2024-11-28 ★ Open-Sora Plan Peking University model 8B
2024-11-27 Spatiotemporal Skip Guidance KAIST paper
2024-11-26 ShowUI NUS model
2024-11-25 Model Context Protocol (MCP) Anthropic library
2024-11-22 ★ Tülu 3 Ai2 model 3.75K2.6K7
2024-11-22 XGrammar CMU library
2024-11-22 ★ Zamba2 (Hybrid SSM/Transformer Suite) Zyphra model 19311037.4B3T
2024-11-22 Amazon Doubles Anthropic Investment to $8 Billion CNBC Amazon news
2024-11-21 ★ AIMv2 Apple paper 1.42K2
2024-11-21 DINO-X 2 IDEA Lab model paper 1.39K4
2024-11-21 Natural Language Reinforcement Learning University College London paper
2024-11-20 Hymba NVIDIA paper 2792
2024-11-20 BALROG University College London eval
2024-11-19 Aquila-VL-2B BAAI model 49
2024-11-19 Loss-to-Loss Prediction Harvard paper
2024-11-18 Pixtral Large Mistral model 124B128K7
2024-11-18 STILL (Slow Thinking with LLMs) Renmin (RUC) paper
2024-11-15 MARS (Make vAriance Reduction Shine) ByteDance paper 721
2024-11-15 LLaVA-CoT Peking University model
2024-11-13 CamemBERT INRIA model 112M
2024-11-12 Spider 2.0 HKU eval 632
2024-11-08 ★ Sarashina2-8x70B SB Intuitions model 12
2024-11-08 SB Intuitions releases 460B-parameter Japanese LLM Sarashina2-8x70B for academia and industry SB Intuitions SB Intuitions news
2024-11-07 SVDQuant + Nunchaku MIT paper
2024-11-07 OneProt TU Munich paper
2024-11-06 Touchstone Johns Hopkins eval
2024-11-04 ★ Hunyuan-Large 2 Tencent model paper 1.59K4826389B52B7T256K
2024-11-04 Hunyuan3D 1.0 Tencent model 3.48K
2024-11-04 TableGPT2 Zhejiang University model 72B
2024-11-01 SimpleQA OpenAI eval 124.33K
2024-11-01 InternThinker PJLab model
2024-11-01 TREC iKAT University of Amsterdam eval
2024-10-31 SelfCodeAlign UIUC paper
2024-10-30 Kinetix University of Oxford paper
2024-10-29 Agentforce Platform Launched for Enterprise AI Agents Salesforce Salesforce news
2024-10-28 Arithmetic Without Algorithms Technion paper
2024-10-28 AutoGLM 2 Z.ai model paper 1
2024-10-28 Zhongguancun Institute of Artificial Intelligence Established ZGCI ZGCA news
2024-10-27 Llama Scope Fudan University paper
2024-10-27 ThunderKittens Stanford library
2024-10-25 SaprotHub Westlake University library
2024-10-24 Infinity-MM BAAI dataset
2024-10-24 MotionCLR 2 IDEA Lab library paper 17
2024-10-24 O1 Replication Journey Shanghai Jiao Tong University paper
2024-10-24 ★ Skywork-Reward 2 Skywork model 37
2024-10-23 DiffuGPT / DiffuLLaMA HKU paper
2024-10-22 OmniGen BAAI model 4.33K171
2024-10-21 Pangea CMU model
2024-10-17 Janus 4 DeepSeek model paper dataset 17.75K10.89K11
2024-10-15 LAPA KAIST model
2024-10-15 OKAMI UT Austin paper
2024-10-15 Zyda-2 Zyphra dataset
2024-10-14 Agent-as-a-Judge KAUST, Meta paper
2024-10-14 Deep Compression Autoencoder MIT model
2024-10-14 TULIP University of Amsterdam model
2024-10-12 OpenR University College London library
2024-10-11 Baichuan-Omni 2 Baichuan paper model 2737B
2024-10-10 MathCoder2 / MathCode-Pile CUHK model
2024-10-10 SIFT (Active Fine-Tuning at Test Time) ETH Zürich paper
2024-10-10 Orthrus Toronto & Vector Institute model
2024-10-10 ★ RDT (Robotics Diffusion Transformer) Tsinghua model 1.2B
2024-10-10 GenARM University of Maryland paper
2024-10-09 REPA — Representation Alignment NYU paper
2024-10-09 MLE-bench OpenAI eval 1.57K975
2024-10-09 ★ PLaMo-100B PFN model 118100B2T
2024-10-09 ★ F5-TTS Shanghai Jiao Tong University model
2024-10-09 HELM Stanford eval
2024-10-09 Demis Hassabis & John Jumper Awarded Nobel Prize in Chemistry Google DeepMind Google news
2024-10-08 LightRAG HKU library
2024-10-07 ★ Falcon Mamba TII model 123.63K37.27B5.8T8.19K
2024-10-03 LLaVA-Critic NTU model
2024-10-03 HELMET Princeton eval
2024-10-03 ProLong Princeton model
2024-10-03 LLMs Know More Than They Show Technion paper
2024-10-03 SageAttention Tsinghua library
2024-10-02 Knowledge Entropy Decay KAIST paper
2024-10-02 GFlowNets Mila paper
2024-10-02 VinePPO Mila paper
2024-10-02 ★ Llama-3.1-Nemotron-70B NVIDIA model 584128K
2024-10-02 Poolside Raises $500M Series B at ~$3B Valuation, Led by Bain Capital Ventures Crunchbase News Poolside news
2024-10-01 RATIONALYST Johns Hopkins paper
2024-10-01 TxT360 MBZUAI dataset 22
2024-10-01 DSPy Stanford library
2024-09-29 Arc2Face Imperial model
2024-09-28 veRL (HybridFlow) HKU library
2024-09-27 ★ Emu3 2 BAAI paper 2.42K4
2024-09-27 Realistic Evaluation of Model Merging Toronto & Vector Institute paper
2024-09-25 ★ Molmo Ai2 model 9132.17K8689/100
2024-09-25 ProX (Programming Every Example) Shanghai Jiao Tong University dataset
2024-09-23 ★ AMPLIFY Mila model 350M
2024-09-23 MobileUI Dataset Xiaomi dataset 79
2024-09-23 MobileVLM Xiaomi model 79
2024-09-19 ★ Qwen2.5 3 Alibaba model paper 27.3K7472B (max)18T8
2024-09-19 Scaling FP8 Training to Trillion-Token LLMs Technion paper
2024-09-18 Qwen2-VL Alibaba model 27.42K76
2024-09-18 Qwen2.5-Coder 2 Alibaba model paper 16.61K37480B (max)7
2024-09-18 Qwen2.5-Math 2 Alibaba model paper 1.08K1272B (max)
2024-09-17 Promptriever Johns Hopkins model
2024-09-15 Berkeley Function-Calling Leaderboard UC Berkeley eval
2024-09-12 ★ o1 OpenAI model ~3.5T200K156/100
2024-09-11 FuXi (weather & ocean FMs) Fudan University model
2024-09-11 ★ Pixtral 12B Mistral model 4.18K612B128K
2024-09-09 Robot Utility Models NYU model
2024-09-05 ★ AdEMAMix Optimizer Apple paper
2024-09-05 ★ DeepSeek-V2.5 DeepSeek model 7.71K
2024-09-05 ★ MiniCPM3-4B OpenBMB model 9.42K40.57K4B
2024-09-05 Open-MAGVIT2 Tencent library 1.01K
2024-09-05 FireRedTTS Xiaohongshu model
2024-09-05 Silvio Savarese Named to TIME 100 Most Influential in AI Salesforce Salesforce news
2024-09-03 ★ OLMoE Ai2 model 1.03K103.29K76.9B1.3B5.13T4.1K
2024-09-01 MLC-LLM / WebLLM CMU library
2024-08-31 Hailuo AI (Video-01 / 2.3) MiniMax model
2024-08-29 CogVLM2 Z.ai model 6
2024-08-28 Auxiliary-Loss-Free Load Balancing Strategy DeepSeek paper 6
2024-08-28 STORM / Co-STORM Stanford library
2024-08-26 Fire-Flyer AI-HPC: Cost-Effective Software-Hardware Co-Design DeepSeek paper
2024-08-22 ★ Show-o NUS model
2024-08-21 Minitron NVIDIA paper 3803
2024-08-21 ★ Sarashina2 SB Intuitions model 91370B2.1T4.1K
2024-08-15 LiveCodeBench UC Berkeley eval
2024-08-13 ★ SWE-bench (Verified / Multimodal / Multilingual) Princeton eval 500
2024-08-12 MovieSum University of Edinburgh dataset
2024-08-12 ★ Tanuki UTokyo model 47B13B1.7T
2024-08-12 CogVideoX: Text-to-Video Diffusion Models Z.ai model 12.78K16
2024-08-08 Exascale Climate Emulator KAUST model
2024-08-07 ★ EXAONE 3.0 LG model 41.67K27.8B
2024-08-06 LLaVA-OneVision NTU model
2024-08-05 MiniCPM-V 2.6 OpenBMB model 25.58K238B
2024-08-04 scTab TU Munich paper
2024-08-02 GPUDrive NYU library
2024-08-01 Chatbot Arena UC Berkeley, Stanford eval
2024-08-01 EXAONEPath 1.0 LG paper 1
2024-08-01 MiniMax Music Series MiniMax model
2024-08-01 KTransformers Tsinghua library
2024-08-01 Pinal / Denovo-Pinal Westlake University model 16B
2024-07-31 Large Language Monkeys Stanford paper
2024-07-29 ★ Apple Foundation Models (AFM) Apple paper 4
2024-07-29 MindSearch PJLab library 6.87K2
2024-07-24 ★ Mistral Large 2 Mistral model 6.26K123B128K8
2024-07-24 Model Collapse University of Oxford paper
2024-07-23 OpenHands CMU library
2024-07-23 ★ Meditron EPFL model 70B (max)
2024-07-23 AbdomenAtlas Johns Hopkins dataset
2024-07-23 ★ Llama 3.1 Meta model 213.13K405B15.6T128K739/100
2024-07-22 RazorAttention: KV Cache Compression Through Retrieval Heads Huawei paper
2024-07-21 Dynamic Memory Compression University of Edinburgh, NVIDIA paper
2024-07-21 ★ Debating with More Persuasive LLMs Leads to More Truthful Answers University College London paper
2024-07-21 Clifford-Steerable CNNs University of Amsterdam paper
2024-07-20 Consent in Crisis: The Rapid Decline of the AI Data Commons Cohere paper 11
2024-07-20 ★ Falcon 2 TII model 4.48K211B5.5T8.19K
2024-07-18 Mistral NeMo Mistral model 32.02K12B128K
2024-07-17 ★ LMMs-Eval NTU library
2024-07-16 Does Refusal Training Generalize to the Past Tense? EPFL paper
2024-07-16 BRIGHT HKU eval 1.38K
2024-07-16 Goldfish / MiniGPT4-video KAUST model
2024-07-16 Codestral Mamba Mistral model 47.98K7.3B
2024-07-15 sbi Tübingen library
2024-07-11 EchoMimicV2 & V3 Ant Group paper 4.25K2
2024-07-11 FlashAttention-3 Princeton paper
2024-07-11 Skywork-Math Skywork paper
2024-07-11 Q-GaLore UT Austin library
2024-07-10 Training on the Test Task Confounds Evaluation Tübingen paper
2024-07-08 Anole Shanghai Jiao Tong University model
2024-07-06 Kolors 2 Kuaishou model paper 4.61K401
2024-07-05 ★ SenseNova 5.5 SenseTime model
2024-07-05 Vimi SenseTime model
2024-07-04 MiniGPT-Med KAUST model
2024-07-04 ★ LLM-jp (v1/v2) NII model 513B
2024-07-04 ★ Step-2 StepFun model 1T
2024-07-03 LivePortrait 2 Kuaishou library paper 18.53K7.37K9
2024-07-03 ★ InternLM2.5 PJLab model 7.22K106.78K1M
2024-07-02 InternVL 2.0 PJLab model 10.06K305.47K108B (max)
2024-07-01 rsl_rl ETH Zürich library
2024-07-01 Copyright Traps for LLMs Imperial paper
2024-07-01 QPlanner LG paper
2024-07-01 Mathstral 7B Mistral model 19.29K7B
2024-07-01 MMLongBench-Doc PJLab dataset 1471
2024-07-01 ★ Agentless UIUC library
2024-06-28 ERNIE 4.0 Turbo Baidu model
2024-06-28 ★ YuLan Renmin (RUC) model 12B1.7T
2024-06-26 Zhang Hongjiang, founder of BAAI: 'AI systems should never be able to deceive humans' Financial Times BAAI news
2024-06-25 HEST-1k & TRIDENT Harvard dataset
2024-06-24 Mooncake 2 Moonshot AI paper dataset 5.55K13
2024-06-24 ★ Cambrian-1 NYU model
2024-06-24 ★ Large Vocabulary Size Improves Large Language Models SB Intuitions paper 1
2024-06-21 NAVSIM Tübingen eval
2024-06-20 ★ Claude 3.5 Sonnet Anthropic model 200K811/100
2024-06-19 AgentDojo ETH Zürich eval
2024-06-19 ★ Semantic Entropy Hallucination Detection University of Oxford paper
2024-06-18 Transformer Neural Processes (TE-TNP / Gridded TNP) University of Cambridge paper
2024-06-18 OlympicArena Shanghai Jiao Tong University eval 11.16K
2024-06-17 AquilaMed-RL BAAI model 241
2024-06-17 DeepSeek-Coder-V2 2 DeepSeek model paper 6.83K4.01K486
2024-06-17 Gaussian Splatting SLAM (MonoGS) Imperial library
2024-06-17 ★ Nemotron-4 340B NVIDIA model 5437340B9T4.1K
2024-06-17 Mip-Splatting Tübingen paper
2024-06-17 DataComp-LM (DCLM) TU Munich dataset
2024-06-17 GaussianAvatars TU Munich paper
2024-06-17 MeshGPT TU Munich paper
2024-06-17 ★ DataComp-LM (DCLM-7B) University of Washington model 7B
2024-06-17 MINT-1T University of Washington dataset
2024-06-17 4K4D Zhejiang University model
2024-06-14 MASt3R Naver paper 2.99K1
2024-06-14 Goldfish Loss University of Maryland paper
2024-06-13 4M-21 EPFL model
2024-06-13 OpenVLA Stanford, UC Berkeley model 7B
2024-06-12 SciRIFF Ai2 dataset 1
2024-06-12 Magpie University of Washington dataset
2024-06-11 Dasheng 3 Xiaomi paper model 4247B
2024-06-10 LlamaGen ByteDance model 1.95K5
2024-06-10 Language Models Resist Alignment Peking University paper
2024-06-10 PowerInfer Shanghai Jiao Tong University library
2024-06-10 Merlin Stanford model
2024-06-10 Apple Intelligence Introduced at WWDC 2024 Apple Apple news
2024-06-09 BiGGen Bench KAIST eval 77
2024-06-09 Equivariant Neural Fields University of Amsterdam paper
2024-06-06 ★ Qwen2 2 Alibaba model paper 27.3K4972B (max)14B (max)6
2024-06-06 ★ Kling 2 Kuaishou model paper
2024-06-06 The Prompt Report University of Maryland paper
2024-06-05 GLM-4V Z.ai model 2.33K
2024-06-04 Seed-TTS 2 ByteDance model paper 1.56K4
2024-06-03 ★ Skywork-MoE Skywork model 1405563146B22B
2024-06-01 agentUniverse Ant Group library 2.27K1
2024-06-01 FlagScale BAAI library 518
2024-06-01 Yuan Embedding Inspur model 3.53K
2024-06-01 InfiniBench KAUST eval
2024-06-01 AtomGen & vector-inference 2 Toronto & Vector Institute model library
2024-05-31 ★ Mamba-2 CMU, Princeton paper
2024-05-30 MotionLLM 2 IDEA Lab model paper 3866
2024-05-29 ★ Codestral Mistral model 13.02K22B32K
2024-05-28 Yuan 2.0-M32 Inspur model 1951.68K140B3.7B
2024-05-28 Knowledge Circuits Zhejiang University paper
2024-05-27 RLAIF-V OpenBMB paper 455223
2024-05-24 The Mosaic Memory of LLMs Imperial paper
2024-05-23 DeepSeek-Prover DeepSeek model 5771186
2024-05-23 SimPO Princeton paper
2024-05-22 ★ Baichuan 4 Baichuan model
2024-05-22 FlashRAG Renmin (RUC) library
2024-05-21 ★ Scaling Monosemanticity Anthropic paper
2024-05-16 ★ Grounding DINO 1.5 2 IDEA Lab model paper 1.12K11
2024-05-16 OceanGPT Zhejiang University model
2024-05-15 ByteFF ByteDance model 84
2024-05-14 Piccolo2 Embedding Model 2 SenseTime model paper 145431
2024-05-14 Hunyuan-DiT 2 Tencent model paper 4.29K11.5B
2024-05-13 AWQ MIT paper
2024-05-13 ★ GPT-4o OpenAI model ~720B128K86/100
2024-05-13 Plot2Code Tencent dataset 241
2024-05-13 RLHFlow (RLHF Workflow) UIUC library
2024-05-08 AlphaFold 3 Google paper
2024-05-07 ★ DeepSeek-V2 2 DeepSeek model paper 5.01K5.35K102236B21B6
2024-05-07 ★ Granite Code IBM model 1.25K1034B4.5T
2024-05-07 VeRA University of Amsterdam paper
2024-05-07 ★ SaProt Westlake University model 1.3B (max)
2024-05-06 SWE-agent Princeton library
2024-05-02 ★ Prometheus / Prometheus 2 KAIST model
2024-05-01 ReFT / pyreft Stanford library
2024-05-01 ProTrek Westlake University model
2024-04-26 llm-jp-corpus NII dataset 474
2024-04-25 InternVL 1.5 PJLab model 10.06K9.81K1626B
2024-04-25 ShareGPT-4o PJLab dataset 16
2024-04-24 ★ SenseNova 5.0 SenseTime model 200K
2024-04-22 OpenELM Apple model 23B
2024-04-22 SEED-X Tencent model 5581
2024-04-18 ★ Reka Core, Flash, and Edge Reka model 125367B128K639/100
2024-04-17 ★ ABAB 6 / 6.5 MiniMax model
2024-04-17 ★ Mixtral 8x22B Mistral model 4.66K141B39B64K
2024-04-16 MiniCheck UT Austin library
2024-04-12 ★ MiniCPM-V 4 OpenBMB model paper 25.58K163.09K239B
2024-04-11 ★ OSWorld HKU eval 369
2024-04-11 MiniCPM-V 2.0 OpenBMB model 67.98K2.8B
2024-04-11 ★ Rho-1: Not All Tokens Are What You Need PJLab, Microsoft paper
2024-04-03 VAR (Visual Autoregressive Modeling) ByteDance model 8.7K7
2024-04-03 PiSSA Peking University paper
2024-04-02 Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks EPFL paper
2024-04-02 ★ HyperCLOVA X Naver model 6
2024-04-01 RULER: What's the Real Context Size of Your LLM? NVIDIA eval 1113
2024-03-28 ★ Jamba AI21 Labs model 51541398B94B256K622/100
2024-03-28 Dataverse Upstage library 564
2024-03-28 sDPO Upstage paper 2
2024-03-28 Jamba: First Production-Grade SSM-Transformer Hybrid Released AI21 Labs AI21 Labs news
2024-03-26 2D Gaussian Splatting Tübingen paper
2024-03-25 PairS University of Cambridge paper
2024-03-23 ★ Step-1 StepFun model 130B
2024-03-23 Step-1V / 1.5V / 2V StepFun model
2024-03-23 Understanding Emergent Abilities from the Loss Perspective Z.ai paper 1
2024-03-22 UniTraj EPFL library
2024-03-19 MergeKit Arcee library 7.13K3
2024-03-19 Dated Data Johns Hopkins paper
2024-03-17 ★ Grok-1 SpaceXAI model 314B78B8.19K6
2024-03-12 ★ Command R / R+ Cohere model 33.71K104B128K5
2024-03-11 ★ Stealing Part of a Production Language Model ETH Zürich paper
2024-03-11 Unraveling the Mystery of Scaling Laws: Part I Meituan paper
2024-03-08 DeepSeek-VL DeepSeek model 4.13K8.76K45
2024-03-08 CogView3 Z.ai model 144
2024-03-04 ★ Claude 3 Anthropic model 200K911/100
2024-03-01 Kimi 2M Moonshot AI model 2M
2024-02-28 WanJuan 2.0 (WanJuan-CC) PJLab dataset
2024-02-27 BioT5+ Microsoft paper 4
2024-02-26 Craftax University of Oxford paper
2024-02-26 ★ scGPT 2 Toronto & Vector Institute model
2024-02-23 MegaScale ByteDance library 24
2024-02-22 MATH-Vision (MATH-V) CUHK eval 3.04K
2024-02-21 SDXL-Lightning ByteDance model 72.67K6
2024-02-21 ★ Gemma Google model 26.93K2257B
2024-02-19 AnyGPT Fudan University model
2024-02-15 ★ Gemini 1.5 Pro Google model 2821M86/100
2024-02-15 SAMformer 2 Huawei model paper 1901
2024-02-15 ★ Sora OpenAI model
2024-02-14 Soft Prompt Threats TU Munich paper
2024-02-07 ★ Moirai Salesforce model 1.52K109.96K32311M
2024-02-06 ★ SenseNova 4.0 SenseTime model
2024-02-05 ★ BGE-M3 2 BAAI model paper 11.8K28.87M52
2024-02-05 DeepSeek-Math 2 DeepSeek model paper 3.32K69
2024-02-04 ★ Qwen1.5 Alibaba model 110B (max)6
2024-02-04 Aligner Peking University paper
2024-02-01 ★ OLMo Ai2 model 6.53K2.4K97B2.46T2.05K
2024-02-01 ★ Aya 101 Cohere model 8.81K1213B
2024-02-01 ★ MiniCPM 3 OpenBMB model paper 9.42K3.73K202B (max)
2024-01-31 ★ Dolma Ai2 dataset 1.51K9
2024-01-30 YOLO-World Tencent paper 6.4K28
2024-01-30 infini-gram University of Washington library
2024-01-29 ★ Baichuan 3 Baichuan model
2024-01-23 ★ InternLM2 2 PJLab model paper 7.22K19.21K27
2024-01-22 Binoculars University of Maryland library
2024-01-20 TFLOP Upstage paper 51
2024-01-19 Depth Anything ByteDance model 8.26K22
2024-01-17 ★ AlphaGeometry Google paper
2024-01-17 ★ GLM-4 Z.ai model 7.07K
2024-01-15 SciGLM / SciInstruct Z.ai paper 4
2024-01-11 ★ DeepSeek-MoE 2 DeepSeek model paper 18.65K1616B2.8B
2024-01-09 Lightning Linear Attention Ant Group paper 2
2024-01-09 Baichuan-NPC Baichuan model
2024-01-04 LLaMA Pro Tencent model 5131.37K1
2024-01-02 ★ EasyEdit Zhejiang University library
2024-01-01 VSAG Ant Group library 482
2024-01-01 FlagAI BAAI library 3.87K
2024-01-01 BenchMARL University of Cambridge, Meta library
2024-01-01 scikit-learn INRIA library
2023-12-28 PanGu-pi 3 Huawei model paper 3.16K27B (max)
2023-12-28 ★ Spike No More: Stabilizing the Pre-training of Large Language Models SB Intuitions paper 2
2023-12-23 ★ SOLAR 10.7B Upstage model 51.95K710.7B6
2023-12-22 GraphCast Google paper 6.67K170
2023-12-21 DUSt3R Naver paper 7.19K3
2023-12-21 ★ InternVL: Scaling up Vision Foundation Models PJLab model 10.06K176B
2023-12-20 ★ Emu2 BAAI model 1.77K5647
2023-12-16 Paloma Ai2 dataset
2023-12-11 ★ Mixtral 8x7B Mistral model 60.57K12046.7B12.9B32K5
2023-12-06 ★ Gemini 1.0 Google model 80932K66/100
2023-12-05 ★ MLX Apple library 26.82K
2023-12-05 Lenna 2 Meituan model paper 871
2023-12-05 ReasonDet Meituan dataset 871
2023-12-04 Magicoder / OSS-Instruct UIUC model
2023-11-29 ★ DeepSeek-LLM 2 DeepSeek model paper 7.03K1.57K8767B5
2023-11-29 GNoME (Materials Discovery) Google paper 1.19K
2023-11-28 ★ Falcon (7B / 40B / 180B) TII model 31115180B3.5T2.05K
2023-11-27 MagicAnimate & Make Pixels Dance ByteDance model 10.91K7
2023-11-27 Yuan 2.0 Inspur model 688102.6B (max)
2023-11-27 UniRepLKNet Tencent model 1.07K34
2023-11-22 T-Rex 3 IDEA Lab model paper 2.68K8
2023-11-20 ★ GPQA: Graduate-Level Google-Proof Q&A Anthropic, Ai2 eval 51021198
2023-11-16 JaxMARL University of Oxford library
2023-11-14 Qwen2-Audio 2 Alibaba model paper 2.08K1.48K277B (max)
2023-11-06 CogVLM Z.ai model 6.74K77
2023-11-02 DeepSeek-Coder 2 DeepSeek model paper 23.66K6.35K10733B (max)
2023-10-30 ★ SEA-LION v1 AI Singapore model 7B980B
2023-10-30 ★ Skywork-13B Skywork model 7291113B3.2T4.1K
2023-10-25 DiQAD Baidu dataset 1
2023-10-19 KwaiYiiMath 2 Kuaishou model paper
2023-10-17 ★ ERNIE 4.0 Baidu model
2023-10-17 ★ BitNet Microsoft paper 63
2023-10-13 VideoCrafter Tencent library 5.06K
2023-10-12 ★ Aquila2 BAAI model 4453634B (max)
2023-10-09 ★ Kimi-v1 Moonshot AI model 200K
2023-10-05 ★ MathCoder PJLab paper 339304
2023-10-04 SEED / SEED-LLaMA Tencent model 64118
2023-10-03 LanguageBind / Video-LLaVA Peking University model
2023-10-01 DALL-E 3 OpenAI model
2023-09-29 ToRA: Tool-Integrated Reasoning Agent Microsoft paper 1.12K21
2023-09-28 PLaMo-13B PFN model 16213B4.1K
2023-09-27 ★ Mistral 7B Mistral model 10.81K479.84K2887.3B32K5
2023-09-26 InternLM-XComposer PJLab model 2.92K13.08K31
2023-09-25 Qwen-Agent Alibaba library 16.51K
2023-09-25 qwen.cpp Alibaba library 627
2023-09-21 ★ PengCheng-Mind PCL model 64200B1.5T
2023-09-12 vLLM (original release) UC Berkeley paper
2023-09-11 NExT-GPT NUS model
2023-09-08 Ant Financial LLM Ant Group model
2023-09-08 CodeFuse Ant Group model
2023-09-08 Fin-Eval Ant Group dataset
2023-09-07 ★ Hunyuan-LLM Tencent model
2023-09-06 ★ Baichuan 2 3 Baichuan paper model 4.1K210.77K12513B2.6T
2023-09-01 OpenSPG & OpenAGL Ant Group library 2.12K
2023-09-01 Asclepius KAIST model
2023-09-01 XTuner PJLab library 5.15K
2023-08-31 ★ ERNIE 3.5 Baidu model
2023-08-31 Belebele Meta eval 1109.8K122
2023-08-30 ★ JAIS MBZUAI model 1.06K2330B1.63T
2023-08-29 DISC-MedLLM Fudan University model 13B
2023-08-29 LongBench Z.ai eval 1.19K104.75K21
2023-08-24 Qwen-VL Alibaba model 136.35K139
2023-08-24 Code Llama Meta model 39870B
2023-08-22 Lagent & AgentLego PJLab library 2.26K
2023-08-21 WanJuan 1.0 Corpus PJLab dataset 8
2023-08-20 ViT-Lens Tencent paper 1901
2023-08-18 KwaiYii Kuaishou model 175B
2023-08-15 ★ Aquila 2 BAAI model paper 4451.06K33B (max)
2023-08-11 MiLM-6B Xiaomi model 4586B
2023-08-08 Baichuan-53B Baichuan model 53B
2023-08-04 Weblab-10B UTokyo model 10B600B
2023-08-04 SoftBank launches an OpenAI for Japan: SB Intuitions, building LLMs and generative AI in Japanese TechCrunch SB Intuitions news
2023-08-03 ★ Qwen 3 Alibaba model paper 21.27K18.12K8972B (max)5
2023-08-02 ★ BGE Text Embeddings BAAI model 11.8K13.9M
2023-08-01 FlagEmbedding & C-MTEB 2 BAAI library dataset 11.8K70
2023-08-01 ★ ABAB 5 / 5.5 MiniMax model
2023-07-31 ToolLLM: Facilitating LLMs to Master 16000+ APIs OpenBMB paper 5.66K69
2023-07-30 SEED-Bench Tencent eval 3645819.24K12
2023-07-27 GCG (Universal Adversarial Attacks) CMU paper
2023-07-20 FLASK KAIST paper
2023-07-19 ★ EXAONE 2.0 LG model
2023-07-18 ★ Llama 2 Meta model 112.62K70B2T4.1K5
2023-07-16 ChatDev 2 OpenBMB paper library 33.36K69
2023-07-13 InternVid PJLab dataset 2.28K32
2023-07-11 ★ Emu BAAI model 1.77K29
2023-07-11 Baichuan-13B Baichuan model 2.93K14.72K13B
2023-07-07 ★ SenseNova 2.0 Upgrade SenseTime model
2023-07-06 ★ InternLM-1.0 PJLab model 7.22K1.68K104B
2023-07-05 PanGu-Weather 2 Huawei model paper 1.36K125
2023-07-01 InternEvo PJLab library 420
2023-07-01 OpenCompass PJLab library 7.08K
2023-06-25 ChatGLM2 / ChatGLM3 Z.ai model 13.68K124.79K6B
2023-06-23 MME (Multimodal Evaluation) BAAI eval 17.87K2.37K14
2023-06-21 LMFlow HKUST library
2023-06-20 ★ Phi-1 ("Textbooks Are All You Need") Microsoft model 991.3B
2023-06-20 UniAD SenseTime, PJLab paper 4.64K1
2023-06-17 LMQL ETH Zürich library
2023-06-15 ★ Baichuan-7B Baichuan model 5.65K148.39K7B1.2T
2023-06-14 WebGLM Z.ai paper 1.6K2
2023-06-12 detrex 2 IDEA Lab library paper 2.29K15
2023-06-07 AlphaDev Google paper
2023-06-05 LIBERO UT Austin eval
2023-06-01 DB-GPT Ant Group library 18.97K
2023-06-01 DLRover Ant Group library 1.66K
2023-06-01 LMDeploy PJLab library 7.9K
2023-06-01 ★ RefinedWeb TII dataset 156
2023-05-31 Let's Verify Step by Step OpenAI paper 2.14K30
2023-05-30 GPT4Tools Tencent paper 77033
2023-05-29 Mix-of-Show Tencent paper 43130
2023-05-27 CPM-Bee OpenBMB model 2.41K310B
2023-05-22 ★ Grouped Query Attention (GQA) Google paper 27
2023-05-20 PengCheng-Nebula PCL announcement
2023-05-17 ★ Ziya LLM 4 IDEA Lab model 4.13K1.83K
2023-05-15 C-Eval HKUST eval 13.95K
2023-05-10 PaLM 2 Google model ~340B5
2023-05-05 Otter (LMM) NTU model
2023-05-04 ★ StarCoder ServiceNow model 22.54K19215.5B1T8.19K
2023-05-02 EvalPlus UIUC eval
2023-04-20 ★ MiniGPT-4 KAUST model
2023-04-20 UltraChat & UltraFeedback OpenBMB dataset 2.86K
2023-04-12 Edit-Friendly DDPM Technion paper
2023-04-11 SenseChat / SenseNova Launch SenseTime model
2023-04-11 SenseMirage SenseTime model 10B (max)
2023-04-10 Stable-DINO 2 IDEA Lab library paper 2425
2023-04-06 Grounded SAM 3 IDEA Lab library paper 17.63K90
2023-04-05 ★ Segment Anything (SAM) Meta model 54.33K538
2023-04-01 BMTools OpenBMB library 2.77K
2023-03-31 A Survey of Large Language Models Renmin (RUC) paper
2023-03-27 EVA-CLIP BAAI model 2.68K80
2023-03-27 Qianfan Platform Baidu announcement
2023-03-20 ★ PanGu-Sigma 2 Huawei model paper 71.1T329B
2023-03-14 OpenSeeD 2 IDEA Lab library paper 7592
2023-03-14 ★ GPT-4 OpenAI model ~666B128K76/100
2023-03-14 ★ ChatGLM-6B Z.ai model 41.05K1.28K1766B
2023-03-09 ★ Grounding DINO 2 IDEA Lab model paper 10.25K2.22M245
2023-02-27 ★ LLaMA Meta model 3.9K65B
2023-02-20 MOSS Fudan University model 16B
2023-01-30 ★ BLIP-2 Salesforce model 581.31K914
2023-01-23 Microsoft Extends Multibillion-Dollar OpenAI Partnership Microsoft Microsoft news
2023-01-01 FlagEvaluation BAAI library 13
2022-12-22 Tune-A-Video Tencent paper 4.37K27
2022-12-19 Instructor HKU model
2022-12-15 ★ Constitutional AI Anthropic paper 306
2022-12-06 InternVideo / InternVideo2 PJLab model 2.28K93
2022-12-05 Painter BAAI model 2.6K10
2022-11-30 ★ Speculative Decoding Google paper 34
2022-11-12 AltCLIP & AltDiffusion BAAI model 3.87K111.27K10
2022-11-10 InternImage 2 PJLab model paper 2.83K41
2022-11-02 Chinese CLIP Alibaba model 5.93K53
2022-11-02 Taiyi 3 IDEA Lab model paper 4192
2022-10-06 ByteTransformer ByteDance library 4791
2022-09-30 CodeGeeX 2 Z.ai model paper 8.79K49
2022-09-21 ★ Whisper OpenAI model 102.39K5.05M1.16K1.55B
2022-09-16 CPM-Ant OpenBMB model 50010B
2022-09-03 TuGraph Ant Group library 1.74K
2022-08-24 ★ GLM-130B 2 Z.ai model paper 7.65K296130B
2022-08-01 COYO-700M Kakao dataset 1.26K
2022-07-22 PanGu-Coder 3 Huawei model paper 3.16K5362.6B (max)
2022-07-11 Orca (continuous batching) SNU paper
2022-07-04 SecretFlow Ant Group library 2.67K
2022-06-24 YOLOv6 2 Meituan library paper 5.89K1.75K
2022-06-06 Mask DINO 2 IDEA Lab library paper 1.54K20
2022-06-01 Vision GNN (ViG) 2 Huawei model paper 4.42K194
2022-05-26 Matryoshka Representation Learning University of Washington paper
2022-05-02 ★ OPT (Open Pre-trained Transformer) Meta model 175B
2022-04-26 CogView2 Z.ai, BAAI model 955121
2022-04-12 ★ Training a Helpful and Harmless Assistant (HH-RLHF) Anthropic paper 365
2022-04-05 ★ PaLM Google model 2.13K540B
2022-03-29 ★ Chinchilla (Compute-Optimal Training) Google paper 663
2022-03-25 ★ CodeGen Salesforce model 5.17K1.81K23516.1B
2022-03-20 Delta Tuning 2 OpenBMB paper library 1.04K
2022-03-07 DINO (DETR) 2 IDEA Lab library paper 2.81K760
2022-03-04 ★ InstructGPT (RLHF) OpenAI paper 4.29K
2022-03-02 DN-DETR 2 IDEA Lab library paper 60555
2022-02-11 BMTrain OpenBMB library 624
2022-02-08 ★ AlphaCode Google paper
2022-02-07 OFA: One For All Alibaba model 2.56K258
2022-01-30 FEDformer 2 Huawei model paper 802541
2022-01-28 ★ Chain-of-Thought Prompting Google paper 4.25K
2022-01-28 DAB-DETR 2 IDEA Lab library paper 579397
2022-01-25 SPIRAL 2 Huawei model paper 6047
2022-01-24 SenseCore AI Infrastructure SenseTime announcement
2022-01-14 DeepSpeed-MoE Microsoft paper 42.49K55
2021-12-31 ERNIE-ViLG Baidu model 3010B
2021-12-01 ★ EXAONE 1.0 LG model 300B
2021-12-01 tFold Tencent model 158
2021-11-30 Donut Naver paper 6.88K5
2021-11-22 Fengshenbang 3 IDEA Lab library model 4.13K3.14K
2021-11-01 KoGPT Kakao model 1.01K176.2B2.05K
2021-10-13 ByteTrack ByteDance library 6.44K105
2021-10-10 Yuan 1.0 Inspur model 58925245B
2021-09-28 DiffVC 2 Huawei model paper 60425
2021-09-10 ★ HyperCLOVA Naver model 4204B560B
2021-09-03 ★ FLAN (Instruction Tuning) Google paper 69
2021-08-10 ★ Codex OpenAI model 3.26K1.43K12B
2021-07-28 Triton OpenAI library 19.41K
2021-07-15 ★ AlphaFold 2 Google paper
2021-07-12 SPLADE Naver paper 2
2021-07-08 OpenDILab / DI-engine 2 SenseTime, PJLab library 3.62K
2021-07-08 OpenPPL (PPLNN) SenseTime library 1.37K
2021-07-07 ★ HumanEval OpenAI eval 3.26K1.43K164
2021-07-05 ★ ERNIE 3.0 & 3.0 Titan Baidu model 195260B
2021-07-01 Meituan Sky Project Meituan announcement
2021-06-28 PFP / Matlantis PFN model
2021-06-24 CPM-2 BAAI model 16415198B
2021-06-17 ★ LoRA (Low-Rank Adaptation) Microsoft paper 2.45K
2021-06-01 OceanBase Ant Group library 10.15K
2021-06-01 ★ Wu Dao 2.0 2 BAAI model paper 1.75T
2021-06-01 Wu Dao Corpora BAAI dataset
2021-05-26 CogView Z.ai, BAAI model 1.8K383
2021-05-13 Grad-TTS 2 Huawei model paper 60443
2021-05-01 Trustworthy AI White Paper Xiaomi paper
2021-04-29 ★ DINO Meta paper
2021-04-26 ★ PanGu-alpha 2 Huawei, PCL model paper 3.16K94200B
2021-04-01 SapBERT University of Cambridge model
2021-03-20 ★ Wu Dao 1.0 BAAI model 11.3B (max)
2021-03-18 ★ GLM (Original) 2 Z.ai model paper 3.51K2852110B
2021-03-18 P-Tuning Z.ai paper 2.08K265
2021-03-01 M6 Series Alibaba model 48
2021-02-27 Transformer in Transformer (TNT) 2 Huawei model paper 4.42K1.01K
2021-02-26 ★ CLIP OpenAI paper 33.74K5.3K
2021-01-11 ★ Switch Transformer Google paper 361
2021-01-05 ★ DALL-E OpenAI model 1.13K12B
2021-01-01 MS-MARCO-CN Baidu dataset 18
2021-01-01 PaddleNLP Baidu library 12.95K
2021-01-01 PaddleSpeech Baidu library 12.61K
2021-01-01 KoBART SK Telecom model 467127
2020-12-07 HEBO 2 Huawei library paper 2.77K19
2020-12-01 CPM-1 BAAI model 1.58K222.6B
2020-10-22 ★ Vision Transformer (ViT) Google paper 12.58K21.58K
2020-10-08 Deformable DETR 2 SenseTime model paper 3.98K12.41K1.87K
2020-09-12 FuxiCTR / BARS 3 Huawei library paper 1.43K12
2020-08-01 Vega Huawei library 848
2020-07-15 PaddleOCR Baidu library 81.73K
2020-06-11 ★ GPT-3 OpenAI model 3.03K175B300B2.05K
2020-06-08 ★ Liquid Time-constant Networks Liquid AI paper 3
2020-04-08 DynaBERT 2 Huawei model paper 3.16K12119
2020-03-28 MindSpore Huawei library 4.69K
2020-03-13 ★ ProGen Salesforce model 702341.2B
2020-03-10 Bolt Huawei library 958
2020-02-13 ★ DeepSpeed Microsoft library 42.49K
2020-01-23 ★ Scaling Laws for Neural Language Models OpenAI paper 1.51K
2020-01-01 Kunlun XPU Baidu announcement
2020-01-01 PaddleDetection Baidu library 14.24K
2020-01-01 PaddleSeg Baidu library 9.34K
2020-01-01 KoGPT2 SK Telecom model 558
2019-11-27 GhostNet 4 Huawei model paper 4.42K406
2019-09-23 TinyBERT 2 Huawei model paper 3.16K169.89K137
2019-09-17 ★ Megatron-LM NVIDIA library 16.66K835
2019-09-01 NEZHA 2 Huawei model paper 3.16K86
2019-08-23 Ascend 910 Series Huawei announcement
2019-08-01 Chainer PFN library 27
2019-07-29 ★ ERNIE 2.0 Baidu model 74
2019-07-25 ★ Optuna PFN library 14.34K
2019-07-19 SUMBT SK Telecom paper 9024
2019-06-20 Alchemy Tencent dataset 11465
2019-06-01 KoBERT SK Telecom model 1.41K18.01K
2019-04-19 ★ ERNIE 1.0 2 Baidu model paper 770
2019-04-03 CRAFT Naver paper 3.38K58
2019-02-14 ★ GPT-2 OpenAI model 24.92K13.36M1.5B1.02K
2018-12-01 ★ JAX Google library 35.79K
2018-10-11 ★ BERT Google model 340M
2018-10-01 OpenMMLab / MMDetection 2 SenseTime, PJLab library paper 32.75K794
2018-06-11 ★ GPT-1 OpenAI model 117M512
2018-06-01 MACE (Mobile AI Compute Engine) Xiaomi library 5.04K
2018-03-16 ApolloScape Baidu dataset 617
2018-03-01 Kata Containers Ant Group library 8.05K
2017-11-24 StarGAN Naver paper 18
2017-11-14 DuReader Baidu dataset 1.18K51
2017-11-02 ★ VQ-VAE Google paper 1.93K
2017-07-26 Xiao AI Xiaomi announcement
2017-07-20 ★ Proximal Policy Optimization (PPO) OpenAI paper
2017-07-05 Apollo Baidu library 26.66K
2017-06-12 ★ Attention Is All You Need (Transformer) Google paper
2017-06-01 Angel ML Tencent library 6.79K
2017-05-03 AI.SG: New National Programme to Catalyse, Synergise and Boost Singapore's AI Capabilities (NRF to Invest Up to S$150M Over Five Years) Ministry of Digital Development and Information AI Singapore news
2017-03-15 DiscoGAN SK Telecom paper 776732
2017-01-18 ★ PyTorch Meta library 100.65K16.19K
2016-09-30 PaddlePaddle Baidu library 23.94K
2016-04-27 OpenAI Gym OpenAI library 37.22K
2016-01-27 ★ AlphaGo Google paper
2015-11-09 ★ TensorFlow Google library 195.63K8.82K
2015-03-09 ★ Distilling the Knowledge in a Neural Network Google paper 13.96K
2015-02-26 ★ DQN (Deep Q-Network) Google paper
2015-02-11 ★ Batch Normalization Google paper 24.38K
2014-12-17 Deep Speech 1 & 2 Baidu paper