Integration and Usage¶
The models that are available via UCloud AI can be consumed in different ways depending on the use case.
Chat Interface¶
From the UCloud AI access portal and model playground page, users can click on
to open a new session on the integrated chat interface.
The chat interface has two main components, a prompt panel (left) and a sidebar (right).
In the prompt users can enter their requests to the AI model in plain text format.
Files can also be added to the chat context by clicking the + symbol below the text field. The chat supports images, audio, video, text documents, PDFs, EPUBs, and Microsoft Office files.
Furthermore, users can select one of the available models by clicking its name and set the reasoning effort the model uses to generate its reply (low, high, max).
In the sidebar, an orange circle
at the bottom indicates that the connection to the backend is being established. A green circle
will be shown to indicate that the model is ready to receive prompts.
Just above, users can see the current thread's total token use and how much of the model's overall context window it uses.
In addition, the sidebar shows the user's chat history, which is saved in /Home/Inference/Chats in My workspace.
A new thread can be created by clicking the
button.
Threads in the chat history can be renamed or deleted by clicking the button next to the given thread name.
The sidebar can be collapsed by clicking the button.
Important
Deletion of a chat thread is permanent and irreversible.
Developer mode¶
All models are configured with default inference parameter values. These can be adjusted by enabling Developer mode using ther toggle switch located above the sidebar:
Configurable settings and options available in Developer mode include:
Settings: Change basic configurations, including model temperature, maximum completion tokens, and the default system prompt.
Advanced settings: Adjust advanced inference parameters, such as presence penalty and frequency penalty.
Curl: Review a ready-to-use curl command pre-populated with the exact inference parameter values and system prompts you have configured.
Usage: View a detailed overview of token consumption for the specific thread.
App Integration¶
UCloud AI models are natively accessible through various UCloud Apps, allowing users to consume them directly inside active jobs. The level of integration varies from application to application. More information can be found on each app's documentation page.
A list of UCloud AI integrated apps includes:
Note
Model integration may be available only on recent versions of the apps.
API Endpoint¶
UCloud AI inference can also be accessed from external clients via an API endpoint. In this case, an API key is required to access the AI models vi an API request.
At the bottom of the model playground page, users can find examples of how to connect clients such as VS Code, OpenCode, and Codex to the endpoint.
Note
Only a limited set of APIs are supported. Using the chat completions API is recommended when possible.
Contents