Mulai sekarangMulai gratis

Quiz Question 2

A new AI service for a real-time chatbot is experiencing inconsistent response times, especially during traffic spikes. The development team suspects that multiple models deployed on the same cluster are competing for GPU resources. Which GKE tool is designed to intelligently route requests to the least-loaded model instance and improve throughput?

Latihan ini merupakan bagian dari kursus

AI Infrastructure: Deployment Types

Lihat Kursus

Latihan interaktif langsung

Ubah teori menjadi aksi dengan salah satu latihan interaktif kami

Mulai latihan