Repository navigation
Conversation
…arn Prediction's error A ground-truth name that is not in the data input failed compilation with None.get. Say which column is missing instead, in the wording MachineLearningScorer already uses. Closes apache#8864 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Automated Reviewer SuggestionsBased on the
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #8918 +/- ##
============================================
- Coverage 92.60% 92.59% -0.01%
+ Complexity 5045 5042 -3
============================================
Files 1252 1252
Lines 53584 53584
Branches 6672 6673 +1
============================================
- Hits 49619 49618 -1
+ Misses 2319 2316 -3
- Partials 1646 1650 +4
*This pull request uses carry forward flags. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
| config | throughput | MB/s | latency | max Δ latest / 7d | |
|---|---|---|---|---|---|
| 🔴 | bs=10 sw=10 sl=64 | 408 | 0.249 | 23,623/32,448/32,448 us | 🔴 -6.0% / 🔴 +139.2% |
| 🔴 | bs=100 sw=10 sl=64 | 878 | 0.536 | 109,409/158,573/158,573 us | 🔴 +24.4% / 🔴 +68.9% |
| ⚪ | bs=1000 sw=10 sl=64 | 1,031 | 0.629 | 966,075/1,078,404/1,078,404 us | ⚪ within ±5% / 🔴 -19.0% |
Baseline details
Latest main 40b507d from same runner
| config | metric | PR | latest main | 7d avg | Δ latest | Δ 7d |
|---|---|---|---|---|---|---|
| bs=10 sw=10 sl=64 | throughput | 408 tuples/sec | 434 tuples/sec | 960.16 tuples/sec | -6.0% | -57.5% |
| bs=10 sw=10 sl=64 | MB/s | 0.249 MB/s | 0.265 MB/s | 0.586 MB/s | -6.0% | -57.5% |
| bs=10 sw=10 sl=64 | p50 | 23,623 us | 22,770 us | 11,059 us | +3.7% | +113.6% |
| bs=10 sw=10 sl=64 | p95 | 32,448 us | 34,003 us | 13,567 us | -4.6% | +139.2% |
| bs=10 sw=10 sl=64 | p99 | 32,448 us | 34,003 us | 16,824 us | -4.6% | +92.9% |
| bs=100 sw=10 sl=64 | throughput | 878 tuples/sec | 913 tuples/sec | 1,235 tuples/sec | -3.8% | -28.9% |
| bs=100 sw=10 sl=64 | MB/s | 0.536 MB/s | 0.557 MB/s | 0.754 MB/s | -3.8% | -28.9% |
| bs=100 sw=10 sl=64 | p50 | 109,409 us | 106,688 us | 87,809 us | +2.6% | +24.6% |
| bs=100 sw=10 sl=64 | p95 | 158,573 us | 127,506 us | 93,907 us | +24.4% | +68.9% |
| bs=100 sw=10 sl=64 | p99 | 158,573 us | 127,506 us | 106,076 us | +24.4% | +49.5% |
| bs=1000 sw=10 sl=64 | throughput | 1,031 tuples/sec | 1,030 tuples/sec | 1,272 tuples/sec | +0.1% | -18.9% |
| bs=1000 sw=10 sl=64 | MB/s | 0.629 MB/s | 0.629 MB/s | 0.776 MB/s | 0.0% | -19.0% |
| bs=1000 sw=10 sl=64 | p50 | 966,075 us | 962,236 us | 863,168 us | +0.4% | +11.9% |
| bs=1000 sw=10 sl=64 | p95 | 1,078,404 us | 1,049,825 us | 907,700 us | +2.7% | +18.8% |
| bs=1000 sw=10 sl=64 | p99 | 1,078,404 us | 1,049,825 us | 933,684 us | +2.7% | +15.5% |
Raw CSV
config_idx,batch_size,schema_width,string_len,num_batches,total_ms,total_tuples,total_bytes,tuples_per_sec,mb_per_sec,lat_p50_us,lat_p95_us,lat_p99_us
0,10,10,64,20,489.80,200,128000,408,0.249,23622.60,32447.73,32447.73
1,100,10,64,20,2277.11,2000,1280000,878,0.536,109409.01,158572.88,158572.88
2,1000,10,64,20,19399.55,20000,12800000,1031,0.629,966074.85,1078404.07,1078404.07
carloea2
left a comment
There was a problem hiding this comment.
One regression reproduced on this head with a passing exact-case control.
| if (groundTruthAttribute != "") { | ||
| resultType = | ||
| inputSchema.attributes.find(attr => attr.getName == groundTruthAttribute).get.getType | ||
| if (!inputSchema.containsAttribute(groundTruthAttribute)) |
There was a problem hiding this comment.
Please keep the exact-case lookup here. With input column label and configured name Label, schema validation now succeeds because containsAttribute ignores case. The generated script then fails at in2df.drop("Label", axis=1) with KeyError. I reproduced this with the generated code; label passes as a control. The native Python filter also compares names exactly.
What changes were proposed in this PR?
When Sklearn Prediction's "Ground Truth Attribute Name to Ignore" names a column that is not in the data input, compilation failed with
java.util.NoSuchElementException: None.get, which does not tell the user what is wrong. It now fails with "Ground Truth column '' is not in the input table", the same wording MachineLearningScorer uses for its columns.Any related issues, documentation, discussions?
Closes #8864
How was this PR tested?
Updated the existing test in
SklearnPredictionOpDescSpecto check the new message.sbt "WorkflowOperator/testOnly *SklearnPredictionOpDescSpec"passes, and so do scalafmt and scalafix.Was this PR authored or co-authored using generative AI tooling?
Generated-by: Claude Code (Claude Opus 5.5)
🤖 Generated with Claude Code