Integration · ML platform
Hugging Face + OrchestrAI
Catalog exported 2026-09-02 · Hugging Face website
Search Hugging Face models, run inference, resolve datasets, and deploy Inference Endpoints from chat.
OrchestrAI exposes 4 Hugging Face operations: 3 are low-risk (read-only or low-impact), and 1 create or modify resources and run only after you confirm the plan.
What teams use it for
ML teams use OrchestrAI to search the Hugging Face Hub for a candidate model, send it a test prompt through the inference API, and deploy it to an Inference Endpoint when it performs well. Model search, inference, and dataset resolution are low risk, while deploying an endpoint is medium risk and runs on request because it creates a new billable resource. There is no operation to pause, scale, or delete an Inference Endpoint, and no model upload, so manage endpoints and repos in the Hub UI.
Every Hugging Face operation, with its risk level
| Operation | What it does | Risk | Step-level approval |
|---|---|---|---|
Download HuggingFace Dataset |
Resolve a HuggingFace dataset for download | Low risk | No |
HuggingFace Inference |
Run inference against a HuggingFace model | Low risk | No |
List HuggingFace Models |
Search/list models on the HuggingFace Hub | Low risk | No |
Deploy HuggingFace Model |
Deploy a model to a HuggingFace Inference Endpoint | Creates resources | No |
Risk tiers come from the catalog: low is read-only or low-impact, medium creates resources and is reversible, high modifies existing resources, destructive may lose data. Every plan that creates or changes resources is shown with its cost estimate and waits for your confirmation. Operations marked with a step-level approval pause again on their own step. Destructive operations require a typed risk phrase.
What you connect
A Hugging Face credential (stored as huggingface).
Connected-service tokens are envelope-encrypted with a per-record key wrapped by a cloud KMS.
Prompts that work
- Search Hugging Face for text-classification models under 500M parameters sorted by downloads
- Run inference on sentence-transformers/all-MiniLM-L6-v2 with the sentence refund my order
- Deploy meta-llama/Llama-3.1-8B-Instruct to an Inference Endpoint on an A10G in us-east-1
Before anything runs
Every mutation shows its plan, cost estimate, and blast radius, then waits for your confirmation. Destructive operations require a typed risk phrase. Credentials are minted per run through OIDC federation and discarded afterward; nothing you create here is invisible later, because every resource lands in the desired-state ledger where drift is detected and can be converged. Details on the security page.
Frequently asked questions
- Can OrchestrAI delete a Hugging Face Inference Endpoint?
- No. Deploying an endpoint is available as a medium-risk operation, but pausing, scaling, and deleting endpoints are not, so use the Hugging Face UI for those.
- Does running Hugging Face inference through OrchestrAI require approval?
- No. Inference is low risk and changes nothing in your infrastructure, so it runs immediately with your Hugging Face token.
- How does OrchestrAI authenticate to Hugging Face?
- You add a Hugging Face credential once in the connections screen. It is envelope-encrypted with a per-record key wrapped by a cloud KMS and is only decrypted inside the run that needs it.
Related integrations
Try it on your own account
Connect your cloud read-only and see your resources, drift, and costs before anything runs. $5 minimum to start. Unused credits refunded in your first 14 days.
Unused credits refunded in your first 14 days.