Red Hat Brings Enterprise AI at Scale Into Focus
21m
Enterprise AI Moves Into Production
Mike Vizard speaks with Tushar Katarki, head of product for Red Hat AI Platforms, about what it takes to run enterprise AI at scale. The discussion starts with a shift many organizations are now facing. AI pilots have proven that the technology can work. The next challenge is running AI securely, efficiently and alongside the rest of the enterprise IT stack.
Katarki says open source models and open source platforms are improving faster than many teams expected. That creates new options for organizations that want more control over cost, data and infrastructure. It also changes how IT leaders should think about AI. The model is only one part of the stack. Enterprises also need governance, monitoring, policy controls and a platform that can support AI in production.
Inference, Open Source Models and Control
The conversation explores why inferencing has become a central part of enterprise AI at scale. Katarki explains how Red Hat has invested in vLLM as an open source inference engine. He compares its role to the Linux kernel. In the same way Linux helped abstract applications from hardware, vLLM helps connect models to different AI accelerators.
Katarki also discusses llm-d and the need for distributed inferencing. Enterprise teams may run many models across many kinds of hardware. Some workloads need low latency. Others need higher throughput. Some models are small and specialized, while others are large and require distributed GPU resources. A platform has to route those workloads intelligently.
AI Gateways and Agent Governance
The episode also looks at model-as-a-service and the rise of AI gateways. Katarki says organizations need more than model access. They need token tracking, rate limits, quota management, chargeback, guardrails and tool-calling controls. Those capabilities become even more important as AI agents move beyond chatbots into longer-running workflows.
For IT teams, the takeaway is clear. Enterprise AI at scale should be managed as a new class of workload. Teams need to think like model service providers and agent service providers. They must deliver useful AI services while keeping cost, security and governance under control.