Rate Limits
API Express enforces rate limits to ensure fair usage and platform stability. This guide explains the limits on each plan, how to monitor your usage, and how to handle 429 errors gracefully.
Overview
Rate limits apply at three levels:
- Requests per second (RPS) — Maximum API calls per second. Applies at burst level.
- Daily quota — Maximum API calls per day.
- Monthly quota — Maximum API calls per month (your plan's included quota).
If you exceed any of these limits, the API returns a 429 Too Many Requests response. Different limits apply to bulk endpoints (up to 10,000 records per request, counted as 1 API call regardless of records).
Limits by Plan
Rate limits scale with your plan. Here's the full breakdown:
| Plan | Requests/sec | Daily limit | Monthly included | Bulk batch size |
|---|---|---|---|---|
| Free | 10 | 1,000 | 1,000 | 100 records |
| Growth | 100 | 50,000 | 100,000 | 1,000 records |
| Enterprise | 1,000+ (custom) | Custom | Unlimited | 10,000 records |
Burst vs sustained rate
The RPS limit is enforced over a rolling 1-second window. You can burst up to your plan's RPS rate, but you can't sustain a higher rate over longer periods — you'll hit the 429 response.
Overage handling
- Free plan — Once you hit 1,000 monthly calls, the API returns
403 QUOTA_EXCEEDEDuntil the next month. - Growth plan — Once you hit 100,000 included calls, additional calls are billed at overage rates. No hard stop.
- Enterprise plan — Unlimited calls. Rate limits are custom to your plan.
Rate Limit Headers
Every API response includes rate limit headers. Monitor these to know how close you are to your limits.
| Header | Description | Example |
|---|---|---|
X-RateLimit-Limit | Your requests-per-second limit | 100 |
X-RateLimit-Remaining | Requests remaining in current window | 87 |
X-RateLimit-Reset | Unix timestamp when the window resets | 1728139843 |
X-Daily-Quota-Limit | Your daily call limit | 50000 |
X-Daily-Quota-Remaining | Calls remaining today | 47823 |
X-Monthly-Quota-Limit | Your monthly included quota | 100000 |
X-Monthly-Quota-Remaining | Calls remaining this month | 84932 |
Reading the headers in Node.js
const response = await fetch(url, { headers });
// Check remaining calls
const remaining = parseInt(response.headers.get('X-RateLimit-Remaining'));
const resetTime = parseInt(response.headers.get('X-RateLimit-Reset'));
if (remaining < 10) {
console.warn(`Approaching rate limit. Resets at ${new Date(resetTime * 1000)}`);
}
Reading the headers in Python
response = requests.get(url, headers=headers)
remaining = int(response.headers.get('X-RateLimit-Remaining', 0))
reset_time = int(response.headers.get('X-RateLimit-Reset', 0))
if remaining < 10:
print(f"Approaching rate limit. Resets at {reset_time}")
Handling 429 Errors
When you exceed a rate limit, the API returns:
{
"status": "error",
"code": "RATE_LIMIT_EXCEEDED",
"message": "Too many requests. Please slow down.",
"http_status": 429,
"retry_after": 2,
"request_id": "req_1a2b3c4d5e6f"
}
The response also includes a Retry-After header telling you how many seconds to wait before retrying:
Retry-After: 2
Automatic retry with backoff
Always implement retry logic that respects the Retry-After header. Here's a recommended pattern:
async function makeRequest(url, options, maxRetries = 3) {
for (let attempt = 0; attempt <= maxRetries; attempt++) {
const response = await fetch(url, options);
if (response.status !== 429) {
return response;
}
// Get retry delay from header, or use exponential backoff
const retryAfter = parseInt(
response.headers.get('Retry-After') || Math.pow(2, attempt)
);
if (attempt === maxRetries) {
throw new Error('Rate limit exceeded after max retries');
}
console.log(`Rate limited. Retrying in ${retryAfter}s...`);
await sleep(retryAfter * 1000);
}
}
Retry-After header — it's calculated by our server to know exactly when you can retry.Rate limit error codes
| Code | Meaning | Retry-After header? |
|---|---|---|
RATE_LIMIT_EXCEEDED | Per-second limit exceeded | Yes (1-60 seconds) |
DAILY_LIMIT_EXCEEDED | Daily quota exceeded | Yes (until midnight IST) |
BULK_LIMIT_EXCEEDED | Batch size over plan limit | No — reduce batch size |
Requesting Higher Limits
If your application needs higher limits than your current plan allows, you have several options.
1. Upgrade your plan
The simplest path. Growth plan includes 100 RPS and 100,000 monthly calls. Enterprise plan has custom limits tailored to your volume.
2. Request a temporary limit increase
For one-off events (product launches, campaign spikes), we can temporarily raise your limits. Email support@apiexpress.in with:
- Your account ID or API key prefix
- Expected peak RPS and duration
- Reason for the increase (launch, campaign, etc.)
We typically respond within 4 business hours and can often grant temporary increases same-day.
3. Spread load across multiple keys
Since rate limits are per key, using multiple API keys is a common way to effectively multiply your limits. This is legitimate as long as it's for legitimate load distribution (not abuse).
// Round-robin across API keys
const keys = [
process.env.API_KEY_1,
process.env.API_KEY_2,
process.env.API_KEY_3,
];
let keyIndex = 0;
function getNextKey() {
const key = keys[keyIndex];
keyIndex = (keyIndex + 1) % keys.length;
return key;
}
// Each call uses the next key in rotation
Best Practices
Monitor your usage continuously
Set up dashboards and alerts to track rate limit headers. Alert when:
- Remaining RPS drops below 20% of your limit
- Daily or monthly quota usage exceeds 80%
- 429 error rate exceeds 0.1% of total requests
Batch requests where possible
Bulk endpoints are counted as 1 API call regardless of records processed (up to your plan's batch size). Use them aggressively:
- 1 bulk verification with 1,000 records = 1 API call (vs 1,000 individual calls)
- 1 bulk recharge with 10,000 records = 1 API call
- 1 bulk fuel query with 1,000 cities = 1 API call
Cache stable data
Not every call needs to be real-time. Cache responses for data that changes infrequently:
- Fuel prices — Cache for 12-24 hours (updates daily)
- Commodity prices — Cache for 1-6 hours (updates multiple times per day)
- Telecom operator data — Cache for 24 hours (rarely changes)
- Weather forecasts — Cache for 15-60 minutes
- Aadhaar/PAN verification — Never cache (regulatory requirement)
Queue non-urgent requests
For background jobs, bulk operations, and analytics, queue requests and process them at a controlled rate rather than bursting. This keeps you under limits without needing to retry.
Use exponential backoff with jitter
When you must retry, use exponential backoff with a small random jitter. This prevents thundering herd problems when multiple clients hit the same limit simultaneously.
// Exponential backoff with jitter
const baseDelay = 1000; // 1 second
const jitter = Math.random() * 500; // up to 500ms extra
const delay = baseDelay * Math.pow(2, attempt) + jitter;
Watch for the warning header
When you're approaching 80% of your daily or monthly quota, we add a X-RateLimit-Warning: true header to responses. Use this to trigger internal alerts before you hit limits.
Related Documentation
- Authentication — API keys and multiple-key strategy
- Error Codes — Full list of 429 error codes
- Sandbox Testing — Test rate limit handling in sandbox
- Pricing — Compare plans and quotas
Talk to Us About Custom Rate Limits
Whether you need a temporary spike for a launch, or a permanent Enterprise plan with custom limits — we'll find the right fit for your workload.