GLM-OCR: Z.ai launches compact OCR model with CogViT and GLM-0.5B encoder-decoder
Original titleGLM-OCR
AISummary
Z.ai has launched GLM-OCR, a compact, high-performance optical character recognition model built on its self-developed CogViT and GLM-0.5B encoder-decoder architecture.
The model uses a dedicated connection layer for cross-modal alignment and CLIP pre-training on billions of image-text pairs for visual semantic understanding and key token extraction. It is designed to stay lightweight for fast inference.
Source: Z.ai Release Notes · docs.z.aiPublished · added here