What are you looking for ?
VergeIO
RAIDON

Arcfra Releases Neutree 1.1 for AI Inference with Native GPU Virtualization and Unified Model Governance

Promoting a Model-as-a-Service dedicated to AI Inference

Blog published by ArcFra July 21, 2026

Arcfra has released Neutree 1.1, a Model-as-a-Service platform for enterprise AI inference, adding native GPU virtualization and new model governance capabilities. The release helps enterprises improve GPU utilization and manage model usage, quotas, access control, and security audits across AI applications.

Native GPU Virtualization for More Efficient Compute Usage
Enterprise AI environments often run different types of models at the same time, including LLMs, OCR, ASR, embedding, and rerank models. Traditional GPU passthrough works well for high-performance LLM inference, but lightweight models or high-concurrency small models can leave memory and compute underused when each workload occupies a full GPU.

Neutree 1.1 adds native vGPU support on top of existing GPU passthrough and logical isolation. Users can enable GPU virtualization as needed and split GPU resources by memory and compute capacity. This allows one GPU to run multiple model instances for workloads such as OCR, speech, embedding, rerank, and multi-service concurrent inference.

For performance-sensitive LLM workloads, users can still use GPU passthrough for using full card resources. This gives teams one platform for both large-model performance assurance and small-model resource reuse.

Neutree also provides a global view of node-level GPU usage. Administrators can see how each node and GPU is being used, then allocate, schedule, and split resources based on actual demand. With hard-isolated resource partitioning, multiple model instances can run on the same GPU without interfering with each other, helping enterprises improve GPU utilization and reduce the cost of model service delivery at scale.

New Model Governance Capabilities
As enterprises connect multiple model services to different business systems, platform teams often face four governance challenges:

  1. Limited usage visibility: Token consumption, request distribution, and cost attribution are not tracked in one place, making resource usage hard to measure
  2. Coarse quota control: Business, apps and API Keys cannot easily receive model resources based on actual needs, which can lead to uneven resource usage
  3. Fragmented access management: API keys are scattered across business systems, while model access scope, rate limits, and concurrency limits lack unified control
  4. Limited traceability: Call time, access source, model, request status, and response details are not fully recorded, making troubleshooting and security audits difficult

Neutree 1.1 extends its model gateway with a unified governance layer for internal and external models. Teams can manage usage statistics, quota management, access control, and security audits without changing existing application calling patterns.

Usage Statistics
Neutree tracks token usage and request activity by API key and model. Administrators can see how different apps and models consume resources, supporting capacity planning, cost allocation, and service optimization.

Quota Management
Administrators can set token quotas for each API key. This prevents one business unit or test workload from consuming too many resources and allows teams to allocate model capacity based on priority and demand.

Access Control
Neutree supports rate limits, concurrency limits, and model access scopes for each API key. This helps enterprises enforce least-privilege access, reduce misuse, and keep shared model services stable.

Access Logs and Security Audit
Neutree records request-level details, including source, API key, model, request status, token usage, throughput, latency, and finish reason. Teams can also retain request and response details for troubleshooting, performance analysis, and security audits.

Building Enterprise AI Inference Infrastructure with Arcfra AECP & Neutree
As enterprise AI moves from pilots to production, model platforms must support deployment, inference, compute management, access governance, and observability.

Neutree 1.1 helps enterprises balance LLM performance with lightweight model resource reuse through native GPU virtualization. Its model governance capabilities bring distributed model calls into one managed entry point, making usage visible, quotas controllable, permissions manageable, and requests traceable.

Neutree is a key component of Arcfra’s agentic AI infrastructure for AI inference. It focuses on compute and model management, model governance, and high-performance inference. Together with Arcfra Enterprise Cloud Platform, it helps enterprises build production-grade, high-performance, and governable model inference infrastructure.

  • Arcfra Kubernetes Engine provides a standard runtime for model inference services, supports GPU driver management, and uses high-performance networking to keep inference instances stable and efficiently scheduled
  • Storage services provide shared and managed storage for model files, reducing duplicate storage and distribution. They can also support high-performance, high-capacity KV cache storage and persistence
  • Observability covers GPU hardware status, container runtime status, inference instance health, model resource usage, access logs, token usage, and request latency. This gives teams end-to-end visibility for troubleshooting, security audit, performance evaluation, and continuous resource optimization

Foxconn, the world’s largest electronics manufacturer, was among the first customers to try Neutree. Foxconn needed to modernize distributed factory infrastructure and improve how AI models were delivered, managed, and observed across environments. By using its existing Arcfra infrastructure as a production-grade AI foundation, Foxconn adopted Neutree to streamline model delivery and unify compute and model resource management. This helped Foxconn accelerate model deployment, improve operational visibility, and build a scalable foundation for future AI and intelligent manufacturing workloads.

Notes
Neutree 1.1 is now open source on GitHub. Visit the project page to learn more about features, deployment, and documentation. You can also submit issues, star the project, or join the community.
GitHub: https://github.com/neutree-ai/neutree.

Read also :
Articles_bottom
AIC