Best practice healthcheck conventions #539
Open
opened 2023-11-30 15:09:31 +00:00 by moritz
·
2 comments
Labels
Clear labels
abra
awaiting-feedback
backups
bug
build
ci/cd
community organising
contributing
coopcloud.tech
design
documentation
duplicate
enhancement
fedi
fedi-infra
finance
funding
good first issue
help wanted
installer
legal
performance
proposal
question
security
test
wontfix
Everything to do with abra
Ping/pong on comms
Something is not working
Go build related issues
Getting the robots into the mix
Opening this thing up
Contributors stuff
Our main website
Design thinking required
Let's write things together
This issue or pull request already exists
New feature
Democratic decision making
Money things
Anything related to grant funding
Easy start with development
Need some help
Installation related issues
Performance related
Large change which requires feedback & decisin making
More information is needed
Securing our shit
Unit or integration test suite
This won't be fixed
No labels
documentation
Milestone
No items
No Milestone
Projects
Clear projects
No projects
Assignees
3wordchant
aadil (Aadil Ayub)
abra-bot (Abra Bot)
ammaratef45
amras (Sarma)
Apfelwurm
BornDeleuze
Brooke
carla
cas (Cassowary)
coopcloud
cyrnel
decentral1se (d1)
dede
devydave
fauno (fauno)
iexos
jade (Jade Ambrose)
jjsfunhouse
jmakdah2 (Jackie Makdah)
joe-irving (Joe Irving)
kawaiipunk (KawaiiPunk)
knoflook
kolaente
lambdabundesverband
linnealovespie (April)
moosemower
moritz
notplants
oxaliq (sorrel)
p4u1
pharaohgraphy (Andrew 🐦🔥❤️🔥✴️)
renovate-bot (Comrade Renovate Bot)
ripclap
simon
sixsmith (Sixsmith)
stevensting
trav (Trav Fryer)
val (val (he/him))
yksflip
Clear assignees
No Assignees
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: toolshed/organising#539
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
This is a very general discussion. I wonder if we could find some good best practice or conventions of how to define good healthchecks in our recipes. This is a very critical topic for us, because in the last time it caused us a lot of nasty problems. Here are the most annoying ones:
start_periodin the recipe again. I would propose to set quite hugestart_periodvalues for each recipe, or are the any arguments against a too highstart_period?I think we should also state this in the docs, as this can cause a lot of pain.
I ran into the same issue on Discourse; if the forum is set to require login, then
GET /serves a 403, instead of a 200 (amusewikidoes similar, but we didn't define a healthcheck for that yet). Solution was to find the/srv/statusendpoint, which works regardless of that setting.Oh yeah that sounds nightmarish. Tuning healthcheck timings is hard; too short and you run into problems like you mention, too long and it increases the chance of walking away from a deployment, not noticing it failed, and then being confused later why the app is still running an old version (or worse, mismatched versions between different services). I wonder if there's a way to make values depend on server load? Otherwise, perhaps a little calculator for different combinations of
interval/retries/timeout/start_periodcould help?The latest changes in
abramake the healthcheck status visible on deployment progress. This might bring back this (necessary) discussion.