Job Type: Contract
Work Mode: Onsite (Client)
Responsibilities
Design build and deliver a Proof of Concept PoC platform for nextgeneration HPCGrid computing solution
Define functional and nonfunctional requirements for enterprise HPC and distributed computing platforms
Architect scalable resilient secure and supportable grid computing solutions
Implement and configure HTCondor preferred or equivalent HPCworkload scheduling platforms
Build and manage distributed compute environments across Windows and Linux platforms
Define platform architecture covering scheduling workload management resource allocation and job lifecycle management
Develop automation solutions for platform deployment configuration and operational management
Create acceptance criteria testing frameworks and validation processes for platform onboarding
Produce onboarding documentation operational runbooks and support procedures
Support onboarding of enterprise application teams including Summit Equities and Global Calypso
Collaborate with CTO Grid teams and consuming application teams to deliver platform capabilities
Manage risks issues dependencies and work tracking for PoC delivery
Ensure solutions meet security auditability performance scalability monitoring and operational readiness requirements
Support platform observability monitoring logging ing and operational support readiness
Influence future strategic direction for grid and distributed computing platforms
To be successful in this role you should have
Strong experience architecting and delivering HPC Grid or Distributed Computing platforms
Deep understanding of workload scheduling concepts including
Queues
Prioritization
Resource Controls
Accounting
Job Lifecycle Management
Handson experience with
HTCondor Preferred
Slurm
GridGain
Apache Ignite
Strong Linux and Windows Server engineering experience
Experience with platform installation configuration troubleshooting and performance tuning
Strong automation and scripting skills using
Python
Bash
PowerShell
Experience with Gitbased development and automation workflows
Experience with Ansible or similar automationconfiguration management tools
Ability to define Functional Requirements FR and NonFunctional Requirements NFR
Experience defining
Availability
Resilience
Scalability
Security
Monitoring
Auditability
Support Models
Strong stakeholder management and delivery management skills
Experience managing plans workboards risks issues actions and dependencies
Experience delivering enterprisescale PoCs and platform implementations
Strong communication skills and ability to work with global engineering teams
Experience working with demanding enterprise stakeholders and application teams
Mandatory Skills : Application Architecture, Ubuntu Linux Administrator