Draft:HippoML
Draft article not currently submitted for review.
This is a draft Articles for creation (AfC) submission. It is not currently pending review. While there are no deadlines, abandoned drafts may be deleted after six months. To edit or make changes to this draft, simply click on the "Edit" tab at the top of the window. To be accepted, a draft should:
It is strongly discouraged to write about either yourself or your business or employer. If you do so, you must declare it. Where to get help
How to improve a draft
You can also browse Wikipedia:Featured articles and Wikipedia:Good articles to find examples of Wikipedia's best writing on topics similar to your proposed article. Improving your odds of a speedy review To improve your odds of a faster review, tag your draft with relevant WikiProject tags using the button below. This will let reviewers know a new draft has been submitted in their area of interest. For instance, if you wrote about a female astronomer, you would want to add the Biography, Astronomy, and Women scientists tags. Editor resources
Last edited by OpalYosutebito (talk | contribs) 4 months ago. (Update) |
| Type | Private (acquired) |
|---|---|
| Industry | Artificial intelligence, Machine learning, Computer software |
| Founded | 2023 |
| Founders | Bing Xu, Hao Lu, Terry Chen |
| Defunct | 2024 |
| Headquarters | Bellevue, Washington, United States |
| Key people | Bing Xu (CEO) Terry Chen (VP of Engineering) Hao Lu |
| Products | HippoEngine, PrivateCanvas |
| Parent | NVIDIA |
HippoML, Inc. was an American artificial intelligence software company that developed inference engines and optimization tools for generative AI models.[1] Founded in 2023 and based in Bellevue, Washington, the company focused on improving deployment efficiency for large language models and other AI workloads across NVIDIA, AMD, and Apple Silicon hardware.[1][2][3] In 2024, HippoML was acquired by NVIDIA.[4]
History
[edit]HippoML was founded in Bellevue, Washington, in 2023 by Bing Xu, Hao Lu, and Terry Chen.[1][5][6][7] Before founding the company, Xu and Lu had worked on deep learning systems including AITemplate, while Chen had also contributed to GPU optimization frameworks including AITemplate.[5][6][7]
During its independent operation, HippoML published engineering results on model serving, attention optimization, quantized inference, and local AI applications.[2][8][9][10][3] In 2024, the company was acquired by NVIDIA.[4]
Technology and products
[edit]HippoML's primary technology was HippoEngine, a GPU inference engine designed to compile machine learning models ahead of time into standalone binary code rather than depend on large runtime environments. HippoML said the engine used a unified runtime and API with support for NVIDIA CUDA, AMD ROCm, and Apple Metal.[2]
HippoML also described a multiple layers model caching system for managing model weights across NVMe SSDs, host memory, and GPU memory. According to the company, this reduced model activation times to about 450 milliseconds.[2]
The company published several blog posts about transformer attention optimization, including variable-length attention, Apple Silicon acceleration, 8-bit HippoAttention, and claims of achieving 1 petaFLOPS attention performance on a single NVIDIA H100 SXM GPU.[8][9][10]
To demonstrate its inference stack on consumer hardware, HippoML released PrivateCanvas, a local desktop AI application for Windows, Linux, and macOS. HippoML described it as an offline tool for image generation and editing that bundled models such as Stable Diffusion XL, Segment Anything, and other generative models into one environment.[3]
HippoML also wrote about decentralized AI inference and the use of local devices for AI workloads in a February 2024 blog post.[11]
Post-acquisition
[edit]Following the acquisition, former HippoML founders and engineers appeared as authors on NVIDIA Technical Blog posts related to AI inference, including posts on automated GPU kernel generation and DeepSeek-R1 inference performance in 2025.[5][6][7][12][13]
See also
[edit]References
[edit]- 1 2 3 "HippoML company profile". PitchBook. Retrieved March 30, 2026.
- 1 2 3 4 "Unified DataCenter & Local Foundation Model Serving: Beyond Docker Way". HippoML Blog. January 8, 2024. Retrieved March 30, 2026.
- 1 2 3 "Super AI Creativity App Run with Local GPU on Mac/Windows/Linux [Early Access]". HippoML Blog. January 2, 2024. Retrieved March 30, 2026.
- 1 2 "Zheng Zhou". Wilson Sonsini. Retrieved March 30, 2026.
- 1 2 3 "Bing Xu". NVIDIA Technical Blog. Retrieved March 30, 2026.
- 1 2 3 "Hao Lu". NVIDIA Technical Blog. Retrieved March 30, 2026.
- 1 2 3 "Terry Chen". NVIDIA Technical Blog. Retrieved March 30, 2026.
- 1 2 "Up to 80X Speedup in Multi-Head Attention on Apple Silicon". HippoML Blog. December 11, 2023. Retrieved March 30, 2026.
- 1 2 "8bit HippoAttention: Up to 3X Faster Compared to FlashAttentionV2". HippoML Blog. January 17, 2024. Retrieved March 30, 2026.
- 1 2 "PetaFLOPS Inference Era: 1 PFLOPS Attention, and Preliminary End-to-End Results". HippoML Blog. February 7, 2024. Retrieved March 30, 2026.
- ↑ "Vision Pro, Decentralized GenAI and AI PC". HippoML Blog. February 7, 2024. Retrieved March 30, 2026.
- ↑ "Automating GPU Kernel Generation with DeepSeek-R1 and Inference Time Scaling". NVIDIA Technical Blog. February 12, 2025. Retrieved March 30, 2026.
- ↑ "NVIDIA Blackwell Delivers World-Record DeepSeek-R1 Inference Performance". NVIDIA Technical Blog. March 18, 2025. Retrieved March 30, 2026.
