Skip to content

LLM-Dev-Ops/benchmark-exchange

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

17 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

LLM Benchmark Exchange

Rust 1.70+

A decentralized platform for standardized LLM benchmarking with community governance

CI License Rust PostgreSQL Redis Docker Kubernetes


Features β€’ Quick Start β€’ CLI β€’ SDK β€’ API β€’ Development β€’ Architecture


Overview

LLM Benchmark Exchange provides a unified platform for the AI community to:

  • Define standardized benchmarks with versioning and metadata
  • Submit model evaluation results with verification
  • Compare models across multiple benchmarks via leaderboards
  • Govern the platform through community proposals and voting
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                         LLM Benchmark Exchange                               β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                                                                              β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚   β”‚  Benchmarks β”‚    β”‚ Submissions β”‚    β”‚ Leaderboardsβ”‚    β”‚ Governance  β”‚  β”‚
β”‚   β”‚  ─────────  β”‚    β”‚  ─────────  β”‚    β”‚  ─────────  β”‚    β”‚  ─────────  β”‚  β”‚
β”‚   β”‚  β€’ MMLU     │───▢│  β€’ Results  │───▢│  β€’ Rankings β”‚    β”‚  β€’ Proposalsβ”‚  β”‚
β”‚   β”‚  β€’ HumanEvalβ”‚    β”‚  β€’ Metrics  β”‚    β”‚  β€’ Compare  β”‚    β”‚  β€’ Voting   β”‚  β”‚
β”‚   β”‚  β€’ Custom   β”‚    β”‚  β€’ Verify   β”‚    β”‚  β€’ Export   β”‚    β”‚  β€’ Comments β”‚  β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚                                                                              β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚   β”‚                         Access Methods                                β”‚  β”‚
β”‚   β”‚   πŸ–₯️  CLI          πŸ“¦  SDK (Rust)         🌐  REST API       gRPC    β”‚  β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚                                                                              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

✨ Features

πŸ“Š Standardized Benchmarks

  • Version-controlled benchmark definitions
  • Rich metadata and documentation
  • Multiple evaluation methods
  • Test case management

πŸ† Leaderboards & Rankings

  • Real-time leaderboard updates
  • Model-to-model comparisons
  • Verification levels
  • Export capabilities

πŸ“ Submissions & Verification

  • Submit evaluation results
  • Multi-level verification
  • Detailed metrics tracking
  • Reproducibility support

πŸ—³οΈ Community Governance

  • Proposal system
  • Democratic voting
  • Transparent decision-making
  • Comment threads

πŸ› οΈ Developer Tools

  • Type-safe Rust SDK
  • Full-featured CLI
  • Shell completions
  • Comprehensive docs

πŸš€ Production Ready

  • Docker & Kubernetes
  • Helm charts
  • CI/CD pipelines
  • Horizontal scaling

πŸš€ Quick Start

Using Docker Compose

# Clone the repository
git clone https://github.com/globalbusinessadvisors/llm-benchmark-exchange.git
cd llm-benchmark-exchange

# Start all services
docker-compose up -d

# The API will be available at http://localhost:8080

Using the CLI

# Install the CLI
cargo install --path crates/cli

# Configure
export LLM_BENCHMARK_API_URL="https://api.llm-benchmark.org"
export LLM_BENCHMARK_TOKEN="your-token"

# Explore benchmarks
llm-benchmark benchmark list
llm-benchmark leaderboard mmlu --limit 10

πŸ“ Project Structure

llm-benchmark-exchange/
β”œβ”€β”€ crates/
β”‚   β”œβ”€β”€ api-rest/        # 🌐 REST API (Axum)
β”‚   β”œβ”€β”€ api-grpc/        # ⚑ gRPC API (Tonic)
β”‚   β”œβ”€β”€ application/     # πŸ“‹ Application services
β”‚   β”œβ”€β”€ cli/             # πŸ–₯️  Command-line interface
β”‚   β”œβ”€β”€ common/          # πŸ”§ Shared utilities
β”‚   β”œβ”€β”€ domain/          # πŸ›οΈ  Core domain types
β”‚   β”œβ”€β”€ infrastructure/  # πŸ—„οΈ  Database & external services
β”‚   β”œβ”€β”€ sdk/             # πŸ“¦ Rust SDK
β”‚   β”œβ”€β”€ testing/         # πŸ§ͺ Test utilities
β”‚   └── worker/          # βš™οΈ  Background jobs
β”œβ”€β”€ migrations/          # πŸ“Š Database migrations
β”œβ”€β”€ docker/              # 🐳 Docker configurations
β”œβ”€β”€ helm/                # ☸️  Helm charts
└── k8s/                 # ☸️  Kubernetes manifests

πŸ–₯️ CLI

The llm-benchmark CLI provides full access to platform functionality.

Installation

cargo install --path crates/cli

Configuration

# Environment variables
export LLM_BENCHMARK_API_URL="https://api.llm-benchmark.org"
export LLM_BENCHMARK_TOKEN="your-api-token"
export LLM_BENCHMARK_OUTPUT_FORMAT="table"

# Or use the config command
llm-benchmark config set api_endpoint https://api.llm-benchmark.org
llm-benchmark config set token your-api-token

Global Options

llm-benchmark [OPTIONS] <COMMAND>

Options:
  -o, --format <FORMAT>    Output format: json, table, plain [default: table]
      --api-url <URL>      API endpoint URL (overrides config)
      --token <TOKEN>      Authentication token (overrides config)
  -v, --verbose            Enable verbose output
      --no-color           Disable colored output
  -h, --help               Print help
  -V, --version            Print version

Commands

πŸ“Š Benchmarks
# List all benchmarks
llm-benchmark benchmark list
llm-benchmark b list                    # Short alias

# List with filters
llm-benchmark benchmark list --category accuracy --status active --limit 20

# Get benchmark details
llm-benchmark benchmark get mmlu

# Create a new benchmark (interactive)
llm-benchmark benchmark create
πŸ“ Submissions
# Submit results
llm-benchmark submit --benchmark mmlu --model gpt-4 --version 0613 --score 0.86

# List submissions
llm-benchmark submit list --benchmark mmlu
πŸ† Leaderboards
# View leaderboard
llm-benchmark leaderboard mmlu
llm-benchmark lb mmlu                   # Short alias

# View with options
llm-benchmark leaderboard mmlu --limit 10 --verified-only
πŸ—³οΈ Governance
# List proposals
llm-benchmark proposal list
llm-benchmark p list                    # Short alias

# View proposal details
llm-benchmark proposal get <proposal-id>

# Create a proposal
llm-benchmark proposal create

# Vote on a proposal
llm-benchmark proposal vote <proposal-id> --approve
llm-benchmark proposal vote <proposal-id> --reject --reason "Needs more detail"

# Comment on a proposal
llm-benchmark proposal comment <proposal-id> "I have a question..."
πŸ” Authentication
# Login
llm-benchmark auth login

# Check authentication status
llm-benchmark auth status

# Logout
llm-benchmark auth logout
🐚 Shell Completions
# Bash
llm-benchmark completions bash > ~/.local/share/bash-completion/completions/llm-benchmark

# Zsh
llm-benchmark completions zsh > ~/.zfunc/_llm-benchmark

# Fish
llm-benchmark completions fish > ~/.config/fish/completions/llm-benchmark.fish

# PowerShell
llm-benchmark completions powershell > llm-benchmark.ps1

πŸ“¦ SDK

The Rust SDK provides type-safe access to the LLM Benchmark Exchange API.

Installation

[dependencies]
llm-benchmark-sdk = { git = "https://github.com/globalbusinessadvisors/llm-benchmark-exchange", package = "llm-benchmark-sdk" }

Quick Start

use llm_benchmark_sdk::{Client, BenchmarkFilter, SdkError};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Create a client
    let client = Client::builder()
        .api_key("your-api-key")
        .build()?;

    // List benchmarks
    let benchmarks = client.benchmarks().list().await?;
    for benchmark in benchmarks.items {
        println!("{}: {}", benchmark.name, benchmark.description);
    }

    // Get a specific benchmark
    let mmlu = client.benchmarks().get("mmlu").await?;
    println!("MMLU has {} submissions", mmlu.submission_count);

    // View leaderboard
    let leaderboard = client.leaderboards().get("mmlu").await?;
    for entry in leaderboard.entries.iter().take(10) {
        println!("#{} {} - {:.1}%", entry.rank, entry.model_name, entry.score * 100.0);
    }

    Ok(())
}

Client Configuration

use llm_benchmark_sdk::Client;
use std::time::Duration;

let client = Client::builder()
    .base_url("https://api.llm-benchmark.org")
    .api_key("your-api-key")
    .timeout(Duration::from_secs(60))
    .retry_count(5)
    .debug(true)
    .build()?;

Environment Variables:

Variable Description
LLM_BENCHMARK_API_URL API endpoint URL
LLM_BENCHMARK_API_KEY Authentication key
LLM_BENCHMARK_TIMEOUT Request timeout (seconds)
LLM_BENCHMARK_DEBUG Enable debug mode

Services

Benchmarks
use llm_benchmark_sdk::{BenchmarkFilter, BenchmarkCategory, BenchmarkStatus};

// List with filters
let filter = BenchmarkFilter::new()
    .category(BenchmarkCategory::Accuracy)
    .status(BenchmarkStatus::Active)
    .page_size(50);

let benchmarks = client.benchmarks().list_with_filter(filter).await?;

// Create a benchmark
let request = CreateBenchmarkRequest::new(
    "My Benchmark",
    "A benchmark for testing X",
    BenchmarkCategory::Accuracy,
);
let benchmark = client.benchmarks().create(request).await?;
Submissions
use llm_benchmark_sdk::{CreateSubmissionRequest, SubmissionResults};
use std::collections::HashMap;

let results = SubmissionResults {
    aggregate_score: 0.95,
    metrics: HashMap::from([("accuracy".to_string(), 0.95)]),
    test_case_results: None,
};

let request = CreateSubmissionRequest {
    benchmark_id: "mmlu".to_string(),
    model_name: "my-model".to_string(),
    model_version: "1.0".to_string(),
    results,
    provider: Some("My Company".to_string()),
    visibility: None,
    notes: None,
};

let submission = client.submissions().create(request).await?;
Leaderboards
// Get full leaderboard
let leaderboard = client.leaderboards().get("mmlu").await?;

// Get top N entries
let top10 = client.leaderboards().top("mmlu", 10).await?;

// Compare two models
let comparison = client.leaderboards()
    .compare("mmlu", "gpt-4", "claude-3")
    .await?;
println!("Score difference: {:.2}%", comparison.score_diff * 100.0);

// Export leaderboard data
let export = client.leaderboards().export("mmlu").await?;
Governance
use llm_benchmark_sdk::{CreateProposalRequest, ProposalType, VoteType};

// Create a proposal
let proposal = client.governance().create(CreateProposalRequest {
    title: "Add new benchmark category".to_string(),
    description: "Proposing to add...".to_string(),
    proposal_type: ProposalType::NewBenchmark,
    rationale: "This would help...".to_string(),
    benchmark_id: None,
}).await?;

// Vote on a proposal
client.governance()
    .vote(&proposal.id.to_string(), VoteType::Approve, Some("Great idea!"))
    .await?;

// Comment on a proposal
client.governance()
    .comment(&proposal.id.to_string(), "I have a question...")
    .await?;

Error Handling

use llm_benchmark_sdk::SdkError;

match client.benchmarks().get("invalid-id").await {
    Ok(benchmark) => println!("Found: {}", benchmark.name),
    Err(SdkError::NotFound { resource_type, resource_id }) => {
        println!("{} '{}' not found", resource_type, resource_id);
    }
    Err(SdkError::Unauthorized { message, .. }) => {
        println!("Auth failed: {}", message);
    }
    Err(SdkError::RateLimited { retry_after }) => {
        println!("Rate limited, retry after {:?}s", retry_after);
    }
    Err(e) if e.is_retryable() => {
        println!("Transient error (already retried): {}", e);
    }
    Err(e) => println!("Other error: {}", e),
}

🌐 API

REST API

The REST API is built with Axum and provides a comprehensive HTTP interface.

Base URL: https://api.llm-benchmark.org/api/v1

Endpoint Method Description
/benchmarks GET List benchmarks
/benchmarks/{id} GET Get benchmark details
/benchmarks POST Create benchmark
/submissions GET List submissions
/submissions POST Create submission
/leaderboards/{benchmark_id} GET Get leaderboard
/proposals GET List proposals
/proposals POST Create proposal
/proposals/{id}/vote POST Vote on proposal

gRPC API

Protocol buffer definitions are available in crates/api-grpc/proto/.


πŸ› οΈ Development

Prerequisites

Requirement Version
Rust 1.70+
PostgreSQL 14+
Redis 7+
Docker 20+ (optional)

Building

# Build all crates
cargo build

# Build in release mode
cargo build --release

# Run tests
cargo test

# Run specific crate tests
cargo test -p llm-benchmark-sdk
cargo test -p llm-benchmark-cli

# Run with all features
cargo build --all-features

Running Locally

# Start dependencies
docker-compose up -d postgres redis

# Run migrations
./migrations/run_migrations.sh

# Start the API server
cargo run -p llm-benchmark-api

# Start the worker (in another terminal)
cargo run -p llm-benchmark-worker

Environment Variables

# Database
DATABASE_URL="postgresql://user:pass@localhost/llm_benchmark"

# Redis
REDIS_URL="redis://localhost:6379"

# API Configuration
API_HOST="0.0.0.0"
API_PORT="8080"
JWT_SECRET="your-secret-key"

# Telemetry
RUST_LOG="info,llm_benchmark=debug"
OTEL_EXPORTER_OTLP_ENDPOINT="http://localhost:4317"

πŸ›οΈ Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                              Presentation Layer                              β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚   REST API  β”‚  β”‚  gRPC API   β”‚  β”‚     CLI     β”‚  β”‚        SDK          β”‚ β”‚
β”‚  β”‚   (Axum)    β”‚  β”‚   (Tonic)   β”‚  β”‚   (Clap)    β”‚  β”‚   (reqwest/async)   β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          β”‚                β”‚                β”‚                    β”‚
          β–Ό                β–Ό                β–Ό                    β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                             Application Layer                                β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”β”‚
β”‚  β”‚                         Application Services                             β”‚β”‚
β”‚  β”‚  β€’ BenchmarkService  β€’ SubmissionService  β€’ GovernanceService           β”‚β”‚
β”‚  β”‚  β€’ UserService       β€’ OrganizationService β€’ ScoringEngine              β”‚β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”β”‚
β”‚  β”‚                              Validation                                  β”‚β”‚
β”‚  β”‚  β€’ Input validation  β€’ Business rules  β€’ Authorization                  β”‚β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                      β”‚
                                      β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                               Domain Layer                                   β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚  Benchmark   β”‚  β”‚  Submission  β”‚  β”‚  Governance  β”‚  β”‚       User       β”‚ β”‚
β”‚  β”‚   Entities   β”‚  β”‚   Entities   β”‚  β”‚   Entities   β”‚  β”‚     Entities     β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”β”‚
β”‚  β”‚                          Domain Events                                   β”‚β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                      β”‚
                                      β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                           Infrastructure Layer                               β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚  PostgreSQL  β”‚  β”‚    Redis     β”‚  β”‚   S3/Minio   β”‚  β”‚    RabbitMQ      β”‚ β”‚
β”‚  β”‚  Repository  β”‚  β”‚    Cache     β”‚  β”‚   Storage    β”‚  β”‚    Messaging     β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Layer Responsibilities

Layer Responsibility Crates
Presentation HTTP/gRPC handlers, CLI commands, SDK api-rest, api-grpc, cli, sdk
Application Use cases, orchestration, validation application
Domain Business entities, rules, events domain
Infrastructure Database, cache, external services infrastructure

🚒 Deployment

Docker

# Build images
docker build -f docker/Dockerfile -t llm-benchmark-api .
docker build -f docker/Dockerfile.worker -t llm-benchmark-worker .

# Run with docker-compose
docker-compose -f docker-compose.prod.yml up -d

Kubernetes

# Apply manifests
kubectl apply -f k8s/namespace.yaml
kubectl apply -f k8s/

# Or use Helm
helm install llm-benchmark ./helm \
  --namespace llm-benchmark \
  --create-namespace \
  --values helm/values.yaml

🀝 Contributing

Contributions are welcome! Please read our contributing guidelines before submitting a PR.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add some amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

πŸ“„ License

This project is licensed under the MIT License - see the LICENSE.md file for details.


Built with ❀️ by the LLM Benchmark Exchange Community

Report Bug β€’ Request Feature β€’ Discussions

About

No description, website, or topics provided.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages