AWS Adds Z.ai’s GLM-5.3 to Bedrock for Eligible Enterprise Customers
The managed service removes the need to operate model-serving infrastructure, while routing profiles, API choice and caching settings shape how teams use it.
Loading page…
The managed service removes the need to operate model-serving infrastructure, while routing profiles, API choice and caching settings shape how teams use it.
Listen to this story
AWS's October 5, 2026, Bedrock release gives eligible enterprises a managed route to Z.ai's GLM-5.3, with US or global cross-Region profiles and no inference infrastructure to provision. The deployment offers OpenAI-compatible and Bedrock-native APIs, prompt-cache controls, and Flex, Standard, and Priority tiers, letting teams choose integration options and balance cost against speed. Z.ai's benchmark figures are company-reported; a current integration limitation remains because Strix's LiteLLM connector does not yet resolve the global profile directly.
Explicit prompt-cache markers require at least 1,024 tokens; AWS recommends them to improve cache hits and reduce input costs and latency.
AWS recommends short-lived credentials for OpenAI-compatible integrations, and accounts need permission to call both the model and its inference profile.
Bedrock charges GLM-5.3 inference per token and retains no persistent resources after requests finish.
Enterprise teams can now use Z.ai’s GLM-5.3 through Amazon Bedrock without operating the infrastructure that runs the model. AWS added the coding and agent model on October 5, 2026, for eligible enterprise customers. Its launch pairs managed APIs with US and Global cross-Region routing, prompt caching and service tiers that trade price against response speed.
GLM-5.3 is designed for complex coding work and extended tasks in which an AI agent reasons through multiple steps and uses tools. AWS positions the Bedrock version as a fully managed alternative to provisioning and operating inference infrastructure—the computing that produces a model’s responses.
Requests use either the us.zai.glm-5.3 or global.zai.glm-5.3 inference profile. A profile is the identifier an application uses to access the model through cross-Region routing. Customers send requests to their chosen source AWS Region, and Bedrock routes them for processing. Eligible users can also try the model without code in the console’s Test > Playground interface.
Authentication is another setup choice. Bedrock can generate API keys for OpenAI-compatible integrations, but AWS strongly recommends short-lived credentials instead of long-lived keys where possible. Its Python example uses the OpenAI software-development kit and aws-bedrock-token-generator to create temporary tokens from existing AWS command-line credentials. The account also needs permissions to call both the base model and the selected inference profile.
Long-running coding agents often resend the same instructions, tool definitions or repository files on each turn. Bedrock enables automatic prompt caching by default, reusing a shared opening portion of the prompt. AWS says this reduces response latency and input-token costs when repeated requests share that context.
Explicit caching gives developers more control: they mark where reusable prompt content ends instead of relying only on automatic detection. Each marked portion must contain at least 1,024 tokens to qualify. AWS recommends this approach and says it can improve cache hits, with corresponding cost and latency savings.
The serving arrangement also changes cleanup. AWS says Bedrock inference is billed per token and leaves no persistent resources, so charges stop when requests finish. Users who generate an API key for the walkthrough should delete it in the Bedrock console when it is no longer needed.
Z.ai reports a 50% improvement over GLM-5.2 on its internal coding benchmark and an 84.5 score on CyberGym. These are company-reported results. AWS says direct comparisons with GLM-5 were not reported because the benchmark tests had changed as the model improved.
AWS’s security-testing walkthrough connects GLM-5.3 to Strix, an open-source agent that searches for vulnerabilities and checks findings with proof-of-concept tests. The target is OWASP Juice Shop, a deliberately vulnerable application running locally. AWS instructs users to test only applications they own or have explicit written permission to assess.
Strix divides the investigation among sub-agents that map the application’s potential attack points, explore vulnerability categories and attempt to validate findings. AWS describes that validation step as a way to reduce false-positive triage. A successful run produces a report with severity, evidence and remediation guidance for each finding.
The example also exposes an integration wrinkle. At publication, LiteLLM—the software Strix uses to connect to model providers—did not yet resolve the global GLM-5.3 profile directly. AWS supplies a workaround that explicitly selects the Converse API route and the inference profile’s full resource identifier.
Loading discussion...
Join the conversation
Explain when faster answers would be worth the extra cost.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.