Distributed AI compute mesh - pool GPUs across machines for shared LLM inference
Infrastructure for creating distributed networks of GPU compute to run LLMs. Pool resources across machines and expose them as a single OpenAI-compatible API endpoint.
Sign in to leave a comment.
No comments yet.