<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>SAIA Platform :: Documentation for AI Services</title><link>https://docs.ai.gwdg.de/en/technical/saia-platform/index.html</link><description>SAIA is the Scalable Artificial Intelligence (AI) Accelerator that hosts our AI services. Such services include Chat AI and CoCo AI, with more to be added soon. SAIA API (application programming interface) keys can be requested and used to access the services from within your code.
API keys are not necessary to use the Chat AI web interface.
The SAIA API is suitable for interactive inference scenarios. If you have a large amount (eg. thousands of LLM queries) of requests that you can process asynchronously, the batch paradigm of our HPC cluster is the better choice. Your batch will be completed more predictably, in less time, and with lower cost. Check out how to get started with our HPC cluster and then running LLMs to learn how you can setup up a batch inference job on the cluster. vLLM is another popular choice for LLM inference.</description><generator>Hugo</generator><language>en</language><atom:link href="https://docs.ai.gwdg.de/en/technical/saia-platform/index.xml" rel="self" type="application/rss+xml"/><item><title>Budget Limits</title><link>https://docs.ai.gwdg.de/en/technical/saia-platform/limits_budget/index.html</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://docs.ai.gwdg.de/en/technical/saia-platform/limits_budget/index.html</guid><description>From a contract to a budget An organisation’s budgets come from what it has bought:
Step Calculation Example Annual budget by contract 72,000 € Monthly budget annual ÷ 12 6,000 € Daily budget monthly ÷ 15 400 € The daily figure is divided by 15, not 30, and that is deliberate. At ÷ 30 you could never use more than an even day’s share, and a deadline or a teaching week would hit the wall. Dividing by 15 gives roughly twice the even daily rate as burst allowance: heavy days are possible, and sustaining that pace all month is not.</description></item><item><title>Rate Limits</title><link>https://docs.ai.gwdg.de/en/technical/saia-platform/limits_rate/index.html</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://docs.ai.gwdg.de/en/technical/saia-platform/limits_rate/index.html</guid><description>How did we calculate rate and budget limits for each organization? Two separate questions, with different answers. An organisation’s limits are derived from its contract, so that what it bought and what it may spend per day agree.
The unit that connects euros to requests A budget is measured in euros and a rate limit in requests, so translating one into the other needs a conversion factor. Derived from real usage statistics at GWDG: An average request consist of roughly 10,000 tokens, which at around 1 € per million tokens comes to about 0.01 € per request.</description></item><item><title>Token cost calculation</title><link>https://docs.ai.gwdg.de/en/technical/saia-platform/token_cost_cal/index.html</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://docs.ai.gwdg.de/en/technical/saia-platform/token_cost_cal/index.html</guid><description>Info The GPU rates and per-token prices below are the figures used to derive the the given examples here. The authoritative and up-to-date prices for each model are published in our current pricing list.
What is our cost? Everything you spend is billed token-exact: the input and output tokens of each request are counted and priced according to the model that served it.</description></item></channel></rss>