Integration and Usage

The models that are available via UCloud AI can be consumed in different ways depending on the use case.

Chat Interface

From the UCloud AI access portal and model playground page, users can click on

to open a new session on the integrated chat interface.

drawing

The chat interface has two main components, a prompt panel (left) and a sidebar (right).

In the prompt users can enter their requests to the AI model in plain text format.

drawing

Files can also be added to the chat context by clicking the + symbol below the text field. The chat supports images, audio, video, text documents, PDFs, EPUBs, and Microsoft Office files.

Furthermore, users can select one of the available models by clicking its name and set the reasoning effort the model uses to generate its reply (low, high, max).

In the sidebar, an orange circle ../_images/icon-yellow-circle.png at the bottom indicates that the connection to the backend is being established. A green circle ../_images/icon-green-circle.png will be shown to indicate that the model is ready to receive prompts.


drawing

Just above, users can see the current thread's total token use and how much of the model's overall context window it uses.

In addition, the sidebar shows the user's chat history, which is saved in /Home/Inference/Chats in My workspace.

A new thread can be created by clicking the

button.

Threads in the chat history can be renamed or deleted by clicking the button next to the given thread name.

The sidebar can be collapsed by clicking the button.

Important

Deletion of a chat thread is permanent and irreversible.

Developer mode

All models are configured with default inference parameter values. These can be adjusted by enabling Developer mode using ther toggle switch located above the sidebar:

drawing

Configurable settings and options available in Developer mode include:

  • Settings: Change basic configurations, including model temperature, maximum completion tokens, and the default system prompt.

  • Advanced settings: Adjust advanced inference parameters, such as presence penalty and frequency penalty.

  • Curl: Review a ready-to-use curl command pre-populated with the exact inference parameter values and system prompts you have configured.

  • Usage: View a detailed overview of token consumption for the specific thread.

App Integration

UCloud AI models are natively accessible through various UCloud Apps, allowing users to consume them directly inside active jobs. The level of integration varies from application to application. More information can be found on each app's documentation page.

A list of UCloud AI integrated apps includes:

Note

Model integration may be available only on recent versions of the apps.

API Endpoint

UCloud AI inference can also be accessed from external clients via an API endpoint. In this case, an API key is required to access the AI models vi an API request.

At the bottom of the model playground page, users can find examples of how to connect clients such as VS Code, OpenCode, and Codex to the endpoint.

Note

Only a limited set of APIs are supported. Using the chat completions API is recommended when possible.