Back to feed
Dev.to
Dev.to
7/16/2026
Inkling MoE + Agent Safety: Token Efficiency Meets Reliability

Inkling MoE + Agent Safety: Token Efficiency Meets Reliability

Short summary

Inkling, a 1T-parameter MoE with 40B active per token and native multimodal I/O, launched on Together Serverless with a reasoning_effort knob for per-request compute control. It collapses multi-model pipelines into one endpoint but self-hosting requires 600GB+ VRAM. Separately, Vercel AI Gateway added Inkling routing and Vercel Connect now generates short-lived OIDC GitHub tokens, replacing long-lived PATs for safer agent credential management.

  • Inkling MoE (1T params, 40B active) offers unified multimodal reasoning with per-request compute tuning on Together Serverless
  • Self-hosting requires Hopper/Blackwell silicon with 600GB+ VRAM; serverless is the practical path for most teams
  • Vercel Connect now mints short-lived OIDC GitHub tokens, eliminating long-lived PATs for agent workflows

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more