We support Palestinian workers.
← Opportunities/Cloudfactory
C
Cloudfactory

Senior Site Reliability Engineer

Berlin, GermanyEnglishAt CloudFactory, we are a mission-driven team passionate…
OperationsEngineeringAI & DataMarketingPartnerships
Want to know how this fits you?

Give Aptiora your real career evidence. We’ll compare it with this opportunity without inventing anything.

Find my fit
At a glance
LocationBerlin, Germany
Work styleNot specified
TypeNot specified
ScheduleNot specified
RemunerationAt CloudFactory, we are a mission-driven team passionate about unlocking the potential…
Start dateAt CloudFactory, we are a mission-driven team passionate about unlocking…
DeadlineOwnership & Drive: Tendency to go above and beyond to…
Required languagesNot stated
The opportunity

About the role

At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale.

Stay in Aptiora while you decide.

Responsibilities, requirements, fit, evidence and preparation are organized here. The original posting stays available for final verification.

Your work

What you’ll do

  • What you’ll own
  • Reliability of platform(includes ML and LLM workloads) - model serving and inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services
  • Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks we run everywhere else
  • Company-wide technical direction: shaping the roadmap and building golden paths that raise the baseline for every team
  • Developer tooling and automation that compounds - reusable GitHub Actions, GitOps workflows, Terraform modules - so every engineer ships faster
  • Reusable components packaging common open-source tools (Grafana, Istio, CloudNative stack, and ML tooling such as model registries and feature stores) for teams to deploy in any environment
  • Secure-by-default infrastructure - baking security, compliance audits, cost governance, and audit trails into the platform in close partnership with our lead/backend/staff engineers.
What matters

What they’re looking for

Select a requirement to inspect fit, evidence or application context.

  • 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes
  • Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm, and Terraform or Cloudformation, on at least one major cloud (AWS preferred).
  • Good proficiency in Python or Go or general scripting for automation and tooling(automation with higher language preferred)
  • First-principles reasoning - Reasoning from constraints and failure modes naming the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off)
  • At least one infrastructure build you owned end to end - with the outcome metric attached (deploy time, MTTR, cost, adoption, availability).
  • Cross-functional strength. Track record working with product, backend/frontend teams to pull through collective initiative.
  • Global Collaboration: Ability to work across global teams and different cultures across various time zones with strong communication skills.
  • Problem Solving: Ability to break down complex problems into simple, actionable solutions.
  • Ownership & Drive: Tendency to go above and beyond to meet deadlines, manage own deliverables, and assist team members.
  • Availability: Willingness to support processes for 24x7 operational support.

Helpful, not always essential

  • Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management
  • Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints(preferred RayServe)
  • MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents)
  • LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request)
  • Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails
Conditions

How they work

  • At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are:
  • Mission-Driven: We focus on creating economic and social impact.
  • People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging.
  • Innovative: We embrace change and find better ways to do things together.
  • Globally Connected: We foster collaboration between diverse cultures and perspectives.
  • If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!

How to apply

  • Ownership & Drive: Tendency to go above and beyond to meet deadlines, manage own deliverables, and assist team members.
  • Apply now and bring your whole, authentic self to work—we can’t wait to meet you!
Conditions

Benefits & working conditions

  • At CloudFactory, we believe that work should be more than just a job—it should be a platform for growth, impact, and community. Here, you’ll earn with purpose, learn every day, and serve a mission that truly matters. If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we’d love to have you on this journey!
  • Join us today and be part of our mission to connect people and technology for a better world! Apply now and bring your whole, authentic self to work—we can’t wait to meet you!
  • Find Jobs in Germany on Arbeitnow
The organization

About Cloudfactory

Verified

We have limited verified information about this company. You can still explore its active opportunities and check the official source.

2active opportunities
1observed locations
AITechnologyEducationManufacturingPublic ImpactStrategyOperationsEngineering
Other live opportunities at CloudfactoryObserved by Aptiora
Forward Deployed EngineerBerlin, Germany
Keep exploring Aptiora

If this one isn’t right.

Related live opportunities you can compare without restarting your search.

Source & verification

Discovery source

Aptiora keeps the source available for trust and final verification while the working experience stays here.

✓ Checked 1 day ago