开始使用免费开始使用

Quiz Question 2

A new AI service for a real-time chatbot is experiencing inconsistent response times, especially during traffic spikes. The development team suspects that multiple models deployed on the same cluster are competing for GPU resources. Which GKE tool is designed to intelligently route requests to the least-loaded model instance and improve throughput?

本练习是课程的一部分

AI Infrastructure: Deployment Types

查看课程

动手互动练习

通过我们的互动练习之一,将理论转化为实践

开始练习