SEARCH · ENGLISH EDITION
Find an AI concept.
Search titles, summaries, article text, source-language titles, and known aliases.
EN
GPU
A highly parallel processor widely used to train and run modern machine-learning models.
EN
Inference
The process of running a trained model on new input to produce predictions, embeddings, or generated output.
EN
Latency and Throughput
Two core serving metrics describing response delay and the amount of work a system completes over time.
EN
Quantization
A technique that represents model weights or activations with lower numerical precision to reduce resource use.