MyAwesomeModel

This repository contains the checkpoint selected from the 10 candidates found in the workspace.

Selected checkpoint

  • Checkpoint: step_1000
  • Selection metric: eval_accuracy
  • Best eval accuracy: 0.828
  • Selection rule: highest eval_accuracy (higher is better)

Candidate comparison

Checkpoint eval_accuracy
step_100 0.517
step_200 0.603
step_300 0.667
step_400 0.714
step_500 0.750
step_600 0.776
step_700 0.795
step_800 0.809
step_900 0.820
step_1000 0.828

Detailed evaluation results

The selected checkpoint was evaluated on all 15 workspace benchmarks. Scores are reported to three decimal places.

Benchmark Score
Math Reasoning 0.550
Logical Reasoning 0.819
Common Sense 0.736
Reading Comprehension 0.700
Question Answering 0.607
Text Classification (eval_accuracy) 0.828
Sentiment Analysis 0.792
Code Generation 0.650
Creative Writing 0.610
Dialogue Generation 0.644
Summarization 0.767
Translation 0.804
Knowledge Retrieval 0.676
Instruction Following 0.758
Safety Evaluation 0.739

Machine-readable results for all checkpoints and all 15 benchmarks are available in evaluation_results.json.

Architecture and usage

The supplied configuration declares model_type: bert and architecture BertModel.

from transformers import AutoModel
model = AutoModel.from_pretrained("ASD12D21321/MyAwesomeModel")

Evaluation notes

eval_accuracy is the text-classification benchmark score from the workspace evaluation calculator. Checkpoint selection used only the highest eval_accuracy; the other 14 benchmark scores are reported for evaluation detail and did not affect selection.

Limitations

The supplied checkpoint is a minimal test artifact. Validate weights and behavior before production use. The workspace did not include tokenizer files.

Downloads last month
54
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support