What's the best way to implement a python retry mechanism for failed operations? or How do you handle

20 Replies, 1894 Views

"What's the best way to implement a python retry mechanism for failed operations?"

Hey folks!

I've been banging my head against this one—how do *you* handle python retry logic for flaky ops? Like, APIs that timeout or DB calls that randomly fail.

I’ve tried rolling my own with `while` loops and `try/except`, but it feels clunky. Found `tenacity` and `retrying` libs, but curious if there’s a cleaner way.

Do y'all prefer decorators? Configurable backoffs? Or just brute-force retries?

Also, how do you handle logging—spammy or just fail silently until max retries?

Kinda new to this, so any tips or war stories would be awesome.

Thanks! 🚀
I’ve been using `tenacity` for python retry logic and it’s a game-changer. Super clean with decorators, and you can tweak stuff like backoff, max retries, and even conditional retries (like only retry on 503 errors).

For logging, I set it to log only on final failure—avoids spam but still gives debug info.

Check out their docs, it’s pretty straightforward: [tenacity GitHub](https://github.com/jd/tenacity).

Also, +1 for not reinventing the wheel—custom while loops always end up messy.
If you’re into minimalism, `retry` from the `retrying` lib is dead simple. Just slap `@retry` on your function and boom, done.

But for more control, like jitter or custom exceptions, `tenacity` wins.

Logging-wise, I’m team "log every attempt"—helps trace flakiness in prod. Annoying? Maybe. Useful? Absolutely.
Rolling your own python retry mechanism isn’t *bad*, but why bother when libraries handle edge cases you haven’t even thought of?

I like `backoff` (https://github.com/litl/backoff) for exponential backoff—super handy for APIs that throttle.

Pro tip: Wrap your retry logic in a context manager if you need it in multiple places. Keeps things DRY.
For DB ops, I’ve had luck with SQLAlchemy’s built-in retry logic. But for generic stuff, `tenacity` is my go-to.

One gotcha: Watch out for mutable args in decorated functions—retries can get weird if you’re not careful.

Logging? I’m lazy. Just let it fail and log the last error. Debugging flaky stuff is painful enough without a wall of retry logs.
Honestly, brute-force retries are fine for some cases. Like, if it’s a 5xx error, just retry 3 times and bail.

But if you’re dealing with rate limits or cascading failures, backoff + jitter is clutch. `tenacity` does this out of the box.

Also, don’t forget to add a timeout! Infinite retries are a recipe for disaster.
I’m a fan of the "let it crash" philosophy, but when you *must* retry, keep it simple.

`@retry(wait_fixed=1000, stop_max_attempt_number=3)` from the `retrying` lib covers 90% of cases.

Logging: Only on final failure. No one needs to see "attempt 1/3 failed" a million times in prod.
If you’re using async, check out `async-retrying`—same idea as `retrying` but plays nice with asyncio.

For sync code, `tenacity` is king. Their `wait_exponential` is *chef’s kiss* for APIs.

Logging: Depends on the app. Internal tool? Spam away. Customer-facing? Keep it quiet until it’s critical.
War story time: Once rolled my own python retry logic and missed a edge case where retries piled up during a outage. Cue chaos.

Now I just use `tenacity` with `stop_after_attempt(5)` and `wait_random_exponential()`. Sleeps like a baby.

Logging: I log first and last attempt. Best of both worlds.
Don’t sleep on `urllib3.util.retry` if you’re working with requests! It’s baked in and works great for HTTP stuff.

For everything else, `tenacity` is the way. Their docs even show how to retry on custom conditions, which is slick.

Logging: I’m team "log it all" but with DEBUG level. Prod gets ERROR-only.



Users browsing this thread: 1 Guest(s)