
Closed
Posted
About Us: We are a leading participant in the Gonka decentralized AI network ([login to view URL]), leveraging high-performance GPU infrastructure to maximize mining rewards. We’re seeking an ML Optimization Engineer to help us achieve superior efficiency and weight in the Gonka ecosystem. Key Responsibilities: - Implement advanced inference optimizations (speculative decoding, quantization, attention modifications, etc.) to maximize mining weight — techniques already proven to achieve double weight with identical GPUs by other participants - Fine-tune Docker configurations for various GPU models based on available registry - Develop custom optimization strategies that balance throughput and quality - Create and maintain custom Docker images optimized for specific GPU architectures - Design and implement systems for stable and scalable mining of Gonka and other protocols - Develop optimized images for Tenstorrent AI ASICs to expand our hardware ecosystem beyond current GPU deployment - Migrate Python code and VLLM implementations to new VLLM images and adapt them for specific GPU cards Required Qualifications: - Proven experience with large language model optimization techniques - Strong understanding of transformer architectures and attention mechanisms - Proficiency with PyTorch, CUDA, and GPU optimization techniques - Experience with vLLM, FlashInfer, or similar inference optimization frameworks - Familiarity with Docker containerization and GPU workload management Preferred Qualifications: - Experience with Claude Code Max (will be provided if needed) - Previous experience with Gonka or similar decentralized AI networks - Background in competitive ML or distributed systems optimization - Experience with NVIDIA GPU architectures (B200/B300/H200/H100/A100) - Knowledge of Tenstorrent AI ASICs or other specialized AI hardware What We Offer: - Opportunity to work with cutting-edge AI infrastructure - High performance-based bonuses tied to achieved weight improvements - Potential for full-time position with percentage of mining profits - Flexible remote work environment - Access to high-end GPU hardware for experimentation Application Process: Intrested candidates should submit their resume along with a brief description of their relevant experience in ML optimization or performance enhancement ideas.
Project ID: 40523838
47 proposals
Remote project
Active 21 secs ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
47 freelancers are bidding on average $52 USD/hour for this job

As an experienced AI and Cloud Developer, I have a versatile skill set that aligns perfectly with your requirements for the ML Optimization Engineer role in the Gonka Network project. My expertise extends from designing backend architectures to building APIs, integrating AI services, and creating responsive web interfaces. This breadth of knowledge enables me to effectively optimize complex systems and balance important factors like throughput and quality, two skills crucial for your project. What sets me apart is my strong proficiency in Python, which is critical given your Python code migration and adaptation tasks. Moreover, I have extensive experience working with PyTorch, CUDA, and GPU optimization techniques - key qualifications you're seeking. I have successfully implemented advanced inference optimizations like speculative decoding, quantization, and attention modifications that can double weight with identical GPUs - a technique you specifically mention.
$50 USD in 40 days
8.3
8.3

Hi — Elias here from Miami. I see you're working on optimizing operations for the Gonka decentralized AI network. The goal here likely revolves around improving efficiency and scalability in your mining operations while ensuring robust performance. A common issue in systems like this is managing the complexities of AI model development and integration with existing infrastructure. What usually matters most here is ensuring that your models can scale effectively without compromising on reliability or maintainability. The tricky part is usually in the coordination between various components, such as data handling and processing pipelines. My approach would focus on creating a modular architecture. This allows for easier updates and integration with new AI models while ensuring that your system remains stable. I have experience working on similar AI optimization projects where I prioritized clean code and seamless integration. A few questions to better understand the scope: Q1 – What user roles will need access to the mining operations and how will permissions be managed? Q2 – Are there specific performance metrics you’re aiming to improve with this optimization? Q3 – What is your current setup regarding backend/API readiness for AI model integration? Happy to go through the details and suggest the best technical approach. Looking forward to hearing from you.
$50 USD in 10 days
7.3
7.3

Hi there, I understand you are looking for an ML Optimization Engineer to improve Gonka mining performance through advanced LLM inference optimization, vLLM tuning, Docker image customization, GPU-specific deployment strategies, and scalable infrastructure across NVIDIA GPUs and future Tenstorrent hardware. I have experience with PyTorch, CUDA, transformer inference optimization, vLLM-based deployments, quantization, attention optimizations, Dockerized AI workloads, GPU performance tuning, and production inference systems focused on maximizing throughput while maintaining output quality. I would analyze the current deployment stack, benchmark bottlenecks, implement GPU-specific optimization techniques, optimize vLLM and container configurations, validate weight improvements through controlled testing, and build a maintainable framework for scaling across different hardware architectures. Q1: Which GPU models are currently deployed in your Gonka infrastructure (H100, H200, A100, B200, etc.)? Q2: What optimization techniques have already been tested, and what performance gains have been achieved so far? Q3: Are there existing custom vLLM forks, Docker images, or benchmarking pipelines that the new engineer will inherit? Best regards.
$50 USD in 40 days
7.2
7.2

Hi I have strong experience with LLM inference optimization, PyTorch, CUDA-aware deployment, vLLM, Docker, GPU workload tuning, quantization, batching, and throughput/latency profiling. The main technical challenge is increasing Gonka mining weight without sacrificing response quality, stability, or compatibility across different GPU architectures. I would approach this by profiling the current container stack, measuring tokens/sec, latency, memory pressure, quality signals, and GPU utilization before applying targeted optimizations. Then I would tune vLLM images, CUDA settings, quantization strategy, speculative decoding, batching, attention backends, and model-serving parameters per GPU type. I can create optimized Docker images for H100, H200, A100, B200/B300, and investigate Tenstorrent-specific deployment paths where the software stack allows it. I can also help migrate Python/vLLM code into newer images, maintain reproducible configs, and document each performance change clearly. My focus would be stable weight improvement through measurable experiments, safe rollback paths, and architecture-specific tuning rather than random parameter changes. I am comfortable working with high-performance AI infrastructure and distributed optimization workflows. Thanks, Hercules
$80 USD in 40 days
6.6
6.6

Hi, I reviewed "ML Optimization Engineer - Gonka Network Mining Operations" and can help with it. The listed skills point to Java Python Matlab and Mathematica Electrical Engineering CUDA Docker AI Model Development AI Development, so I would keep the implementation aligned with that. I would first check the current setup, then complete the work in a way that is easy to review. My focus would be turning the design direction into something polished and usable. Before starting, I would confirm: 1. Where should the completed work be deployed, and what hosting access will be provided? 2. Is there any current codebase, admin panel, or documentation I should review first? 3. Are there specific security, hosting, or maintenance requirements I should keep in mind? Best regards, Houssame
$50 USD in 40 days
6.5
6.5

As a seasoned professional with 8 years of experience as a Full-Stack and AI Engineer, I bring a wealth of knowledge and expertise relevant to the ML Optimization Engineer position. My core focus on LLM-based systems, such as your project entails, coupled with my experience in transformative optimization strategies, uniquely positions me to maximize your mining rewards. I am not only familiar with Docker containerization and GPU workload management, but also have hands-on experience with PyTorch, CUDA, and GPU optimization techniques - essential skills for realizing the weight improvements you're aiming for. My track record of excelling in producing tangible business results through my work is also aligned with what you're seeking. Clients across diverse industries including FinTech, Healthcare, E-Commerce, Travel, and EdTech have benefited from the clean architecture, system reliability, and measurable business impact I foster in my projects - all qualities that I can bring when working on your Gonka Network Mining Operations. Overall, by choosing me, you'll be getting an experienced ML engineer who understands your needs, is familiar with similar decentralized AI networks like Gonka, and can deliver significant improvements to maximize performance in your Mining operations.
$50 USD in 40 days
5.1
5.1

The real bottleneck here isn’t “more GPUs” — it’s mismatched inference stacks and one-size-fits-all images that leave per-card performance on the table. To materially increase Gonka weight you need profile-driven, per-architecture optimizations (speculative decoding, precision reduction, attention kernels) tied to strict quality/throughput targets. I’d start with a short profiling sprint: collect per-GPU latency, memory, and weight impact; rank optimizations by ROI; implement incremental changes (quantization + calibration, FlashAttention/FusedKernels, speculative decoding in vLLM, batch/pipeline tuning) and validate against mining weight metrics. Parallel work: build per-GPU Docker images and a vLLM migration path with automated benchmarks. Suggested stack: PyTorch, vLLM, FlashAttention/FusedKernels, CUDA/cuDNN, TensorRT where useful, ONNX for ASIC paths, NVIDIA Container Toolkit + Docker, and Tenstorrent’s SDK/runtime for ASIC images. Make images CI-driven, benchmark-as-code, and feature-flagged so you can roll back or A/B test optimizations. I led a similar deployment (CrowdAxis): containerized an ML scoring endpoint, introduced inference optimizations and a CI benchmark harness, reducing latency and stabilizing throughput under load. If you want, I can scope a 2-week profiling + optimization roadmap. Quick question: can you share current baseline weight and per-GPU latency/throughput metrics plus the exact GPU/ASIC models in your registry?
$50 USD in 7 days
4.8
4.8

As an AI Development expert and Python specialist with a demonstrated understanding of CUDA and Docker, I believe I am the ideal candidate for your ML Optimization Engineer role. My proficiency in PyTorch aligns perfectly with your project's necessity to implement advanced inference optimizations and fine-tune Docker configurations. Additionally, my experience with vLLM and FlashInfer - frameworks you mentioned, set me apart from other candidates. I’m also no stranger to competitive ML or distributed systems optimization since I have previously engaged in similar projects that have required it. One of the reasons I think I'm a great fit for the job is because of my skill in understanding transformer architectures and attention mechanisms. This would enable me to develop custom optimization strategies for your specific needs, striking a perfect and advantageous balance between throughput and quality.
$50 USD in 40 days
4.2
4.2

Hello, We would like to grab this opportunity and will work till you get 100% satisfied with our work. We are an expert team which have many years of experience on Java, Python, Matlab and Mathematica, Electrical Engineering, CUDA, Docker, AI Model Development, AI Development Lets connect in chat so that We discuss further. Thank You
$60 USD in 40 days
4.4
4.4

Hi, My proposal centers on applying advanced inference optimizations—such as speculative decoding, quantization, and attention modifications—to demonstrably double your mining weight on the Gonka network, leveraging deep expertise in PyTorch, CUDA, and vLLM for maximum throughput. I will systematically fine-tune Docker configurations and custom images for your specific GPU architectures (H100/A100, etc.) to ensure stable, high‑performance operations while balancing output quality. Beyond GPU optimization, I will design scalable systems and migrate your existing codebase to new vLLM images, extending your hardware ecosystem by developing optimized images for Tenstorrent AI ASICs. My approach prioritizes both immediate weight improvements and long‑term reliability, enabling you to outpace competitors and capitalize on performance‑based bonuses. Kind regards Mojjammil
$50 USD in 40 days
3.8
3.8

Hello, I have hands-on experience with projects having similar features. I have 9+ years of experience in web and mobile application development, I can confidently deliver a scalable, user-friendly, and high-performance solution exactly as per your requirements. I have worked on similar platforms with modern UI/UX, API integrations, automation flows, admin management, and secure backend architecture. INCLUDED SERVICES: ✔ UNLIMITED REVISIONS UNTIL SATISFACTION ✔ 3 YEARS OF POST-LAUNCH SUPPORT ✔ FULL SOURCE CODE UPON DELIVERY I focus on quality work, clear communication, and on-time delivery to ensure the project runs smoothly from start to finish. Let’s discuss in chat as I have some queries to ask regarding the project to proceed further. Thanks InvokeTech
$50 USD in 40 days
4.6
4.6

Hi, I reviewed your Gonka ML optimization role carefully. This looks like a strong fit for someone who can work close to the inference stack: Docker, vLLM, PyTorch, CUDA-aware tuning, GPU-specific configuration, quantization, speculative decoding, batching, memory use, and production stability. I would start by benchmarking the current mining setup across your available GPUs, reviewing Docker images, vLLM configuration, model serving parameters, logs, quality metrics, and weight calculation behavior. From there, I would tune inference settings, container builds, batching, quantization options, and GPU-specific profiles while keeping output quality stable. The main thing here is measurable optimization. I would document every change, compare before/after performance, and avoid risky changes that inflate throughput but reduce quality or stability. I have experience with Python, PyTorch, Dockerized AI workloads, LLM serving workflows, performance tuning, GPU infrastructure, inference optimization, and reproducible benchmarking.
$80 USD in 40 days
0.2
0.2

I can help optimize Gonka decentralized AI mining by improving LLM inference efficiency and stability on your GPU fleet. Key focus areas: - Inference optimizations: speculative decoding, quantization strategies, and attention modifications to maximize “mining weight” without sacrificing output quality. - vLLM/Docker pipeline: migrate and adapt existing Python/vLLM implementations into GPU-specific vLLM images; tune Docker workloads across NVIDIA architectures (e.g., H100/A100 class). - Throughput-quality control: build repeatable benchmarking and regression checks so changes improve weight while maintaining reliability. - Scalable mining systems: containerized orchestration and monitoring for stable distributed execution. - Tenstorrent extension: create optimized Tenstorrent AI ASIC images/adapters to expand beyond GPUs. Approach: 1) Baseline current vLLM behavior and measure weight/latency/throughput. 2) Apply targeted optimizations and iterate with controlled experiments. 3) Package results into maintainable GPU-arch-specific Docker images. Efficient, measurable improvements are the goal, because in Gonka, small inference gains can compound.
$50 USD in 29 days
0.0
0.0

Optimizing inference pipelines for Gonka mining rewards requires balancing throughput and GPU utilization through speculative decoding and quantization, precisely what I engineered in production for a 12-agent trading pipeline handling structured BUY/SELL outputs under latency constraints. The solution involves Dockerizing vLLM with architecture-specific optimizations (NVIDIA B200/B300/H200/H100/A100), implementing attention slicing for variable batch sizes, and validating weight improvements against mining benchmarks, same approach that doubled effective capacity on our trading agents while cutting Azure OpenAI costs 38%. Let's start with a one-day spike: containerized inference test on your hardware with real payloads, delivering throughput/weight metrics and GPU utilization plots.
$50 USD in 7 days
0.0
0.0

Hi, This project aligns closely with my experience in AI infrastructure, model serving, GPU optimization, Dockerized deployments, and performance-focused backend systems. I have worked with Python, PyTorch, CUDA-based workloads, inference pipelines, containerized AI environments, and high-throughput distributed architectures where maximizing hardware utilization is critical. For Gonka, my approach would be to systematically benchmark the existing deployment stack, identify bottlenecks across model serving, memory utilization, KV-cache handling, attention implementations, batching strategies, and container configurations, then apply targeted optimizations. I am comfortable working with vLLM-based deployments, custom Docker images, GPU-specific tuning, and migration efforts across different hardware profiles while maintaining stability and reproducibility. I am particularly interested in the challenge of improving mining weight through inference optimization rather than simply scaling hardware. My focus would be on measurable performance gains, repeatable deployment processes, monitoring, and hardware-specific optimization strategies that can be applied across your GPU fleet and future accelerator platforms. I am available to start immediately and can contribute both hands-on engineering work and performance experimentation. Best, Justin
$50 USD in 40 days
0.0
0.0

Washington, United States
Member since Apr 22, 2026
$25-50 USD / hour
$15-25 USD / hour
$25-50 USD / hour
$10-13 USD
₹600-1500 INR
₹600-1500 INR
₹1500-12500 INR
₹12500-37500 INR
€18-36 EUR / hour
$15-25 USD / hour
$250-750 USD
$250-750 USD
₹4000 INR
$2-8 USD / hour
₹12500-37500 INR
$10-30 AUD
₹1500-12500 INR
$250-750 USD
$15-25 USD / hour
$3-10 SGD / hour
$30-250 USD