Follow us
Breaking
Tech News

Modular expands the Mojo ecosystem for AI developers

Modular is scaling the Mojo programming language to bridge the gap between Python and hardware acceleration. Founded by Chris Lattner, the company aims to deliver performance gains of up to 35,000x over Python using MLIR and the LLVM toolchain.

Share

Lattner leads the Mojo 1.0 expansion

Modular reached a milestone this year with the release of Mojo 1.0. Chris Lattner and Tim Davis founded the Palo Alto-based company Modular in 2022 to solve the complexity and fragmentation in AI technical infrastructure by creating a platform that scales machine learning models on both CPUs and GPUs. This company raised $100 million in a funding round led by General Catalyst, with participation from GV, SV Angel, Greylock, and Factory, bringing its total raised funds to $130 million. Lattner, who previously worked at Google and co-founded the LLVM Compiler infrastructure project, the Clang C++ compiler, and the Swift programming language, intends to use these funds to support team growth and product expansion. The Mojo community grew to more than 120,000 developers in the four months following the product keynote in early May. Currently, 30,000 developers remain on the waitlist for the language. Lattner aims to make machine learning and its infrastructure more accessible to non-experts through user-friendly syntax.

Mojo programming specifications

Mojo combines the simplicity of Python with the speed and memory security of Rust. It delivers performance gains of 35,000x over Python. This surpasses the 22x speed improvement of PyPy and the 5,000x improvement of Scalar C++. The language uses Multi-Level Intermediate Representation (MLIR) to scale across hardware types. It compiles into machine code via the LLVM toolchain. Mojo handles variables with let for immutable values and var for mutable values. These restrictions undergo enforcement during compilation to prevent mutation of immutable references. The language uses struct to define types with fixed arrangements for native machine speed, similar to C++ or Rust. The fn keyword defines functions that require explicit typing and local variable declarations. The Mojo SDK includes the mojo driver, a Visual Studio Code extension, and a Jupyter kernel for MacOS and Linux. The mojo driver provides a shell and enables building Mojo applications, packaging modules, generating documentation, and formatting code. The Visual Studio Code extension contains productivity features for developers. A Windows release will arrive soon.

Mojo belongs to the Python family and functions as a superset of Python, allowing users to leverage the entire Python ecosystem. The @struct decorator generates __init__, __copyinit__, and __moveinit__ methods automatically. The @value decorator works on types with copyable or movable members. Mojo’s String is an alias for DynamicVector[SIMD[DType.si8, 1]]. The language includes the DType base construct and SIMD types. The DynamicVector container holds sequences of these types. Mojo performs as fast as Rust on Mac. In one benchmark for the Fibonacci sequence, Python 3 reached a mean time of 16,374.7 us, while the Mojo iteration reached 43,852.7 us. In another test of the Mandelbrot set, the Mojo version with no optimization reached a mean time of 135,880.5 us, while the Mojo parallelize version reached 7,139.4 us.

Feature Mojo Specification
Variable Mutability let (immutable) and var (mutable)
Type Definition struct
Function Definition fn
Primary Compiler LLVM
Acceleration MLIR

Hardware acceleration and the NVIDIA competition

NVIDIA holds a market capitalization of $1 trillion. Developers use the NVIDIA cuTile Python tutorial to build tiled GPU kernels for vector and matrix addition. This workflow requires an NVIDIA driver R580+ or CUDA Toolkit 13.1+. If the system lacks these requirements, the notebook falls back to native PyTorch operations. I find the automated fallback to PyTorch a disappointing solution for teams requiring consistent hardware control. Mojo allows developers to harness the Python ecosystem while optimizing for heterogeneous systems. In one test for the Mandelbrot set, the Mojo parallelize version reached 7,139.4 us, while the Python version reached 5,444,155.4 us. The NVIDIA tutorial shows that cuTile runtimes outperform equivalent PyTorch operations for large workloads such as a 777 x 1001 matrix or a 512 x 768 x 384 tensor. These tiled kernels enable fine-grained parallelism and reduced memory traffic for tasks like a 1,000,003-element vector addition. Can teams truly abandon the stability of existing frameworks for these specialized kernels?

Share

Technewsdaily

Senior tech writer covering AI, gadgets and cybersecurity. Breaking down the news that matters, every day.