Lets continue from our test and try and get something working-ish
Under certain circumstances it seems like Spark's optimistic filter pushdown can lead to double evaluation, for inexpensive filters (like isNull) the increase in evaluation cost is probably ok when combined with the filter ordering improvement in 3.5, but for expensive filters (regular expressions, UDFS, etc) this increase in evaluation time is not ideal.