Researchers and companies are continuing to build and deploy foundation models, a category of AI neural networks trained on large amounts of raw data that can be adapted to a wide range of tasks. According to the 2024 AI Index report from the Stanford Institute for Human-Centered Artificial Intelligence, 149 foundation models were published in 2023, more than double the number released in 2022.
Foundation models generally learn from unlabeled datasets, which reduces the time and expense of manual labeling. The source text says they can be fine-tuned for tasks ranging from translating text to analyzing medical images and performing agent-based behaviors.
The article traces the rise of the field through transformer models, large language models (LLMs), vision language models (VLMs) and other neural networks. It notes that the 2017 paper on transformers helped inspire BERT and other LLMs, while OpenAI’s GPT-3 in 2020 and ChatGPT later pushed broader public use of the technology.
It also says foundation models are becoming multimodal, meaning they can process and generate text, images, audio and video. One example cited is Cosmos Nemotron 34B, described as a leading VLM trained on 355,000 videos and 2.8 million images that can query and summarize images and video from the physical or virtual world.
Another area highlighted is physical AI, which the source describes as enabling autonomous machines like robots and self-driving cars to interact with the real world. It says world foundation models can simulate real-world environments and predict outcomes based on text, image or video input. The text adds that NVIDIA Cosmos world foundation models are trained on 20 million hours of driving and robotics data and are used with the NVIDIA Omniverse platform to generate synthetic data.
The source also says businesses are increasingly customizing pretrained foundation models instead of building from scratch. It cites NVIDIA NeMo framework for building custom chatbots and personal assistants, and mentions retrieval-augmented generation, or RAG, as a method that lets models use external resources like a corporate knowledge base.
The article closes by noting several risks tied to foundation and generative AI models, including:
- amplifying bias in training data,
- introducing inaccurate or misleading information in images or videos, and
- violating intellectual property rights of existing works.
It says current safeguards include filtering prompts and outputs, recalibrating models on the fly and scrubbing large datasets.
Source: blogs.nvidia.com.
Companies can share verified announcements through Newz9’s international press release submission page.

