


Simple AI Inference is a serverless service that does not require infrastructure construction, allowing multiple types of AI models to be called through a single API. In addition, various latest open-source models can be used immediately after application, greatly improving the convenience of development.
Limit Tokens Per Minute(TPM) and Requests Per Minute(RPM) to maintain consistent service availability in the event of a sudden traffic surge. It keeps services stable by precisely controlling data traffic even when processing large amounts of data.
Customer data is never used for model training, ensuring it remains safe from leaks and strictly protected within Samsung Cloud Platform’s secure environment.
It is charged based on the actual number of Input/Cached/Output tokens used. This allows flexible and highly efficient cost management tailored to your service scale and traffic.


Whether you’re looking for a specific business solution or just need some questions answered, we’re here to help