Repository navigation
Selection using Base functions and possibly missing values #134
Description
Activity
I'll have to investigate why the . lifting doesn't work in that context. I'm also going to add a lifted version for the
Dateconstructor in the DataValues.jl package that should make your second attempt work.In the meantime, you can always handle the lifting manually by hand:
@select {i.idx, date = isnull(i.date) ? ?Date() : DataValue{Date}(Date(get(i.date), "mm/dd/yyyy"))}
Yes, cumbersome, but you don't have to wait for me to finish these fixes ;)
Also, if you use the new FileIO.jl integration for loading files that I showed in my juliacon talk, you get a proper
Datecolumn in theDataFramewithout any manipulation, the TextParse.jl package seems to autodetect the type of the column. To get started with that approach, do aPkg.clone("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/davidanthoff/Dataverse.jl.git"), thenusing Dataverse. After that the following should work:df = DataFrame(load("filename.csv"))
or alternatively you can use a pipe syntax:
df = load("filensame.csv") |> DataFrame
Thanks for taking a look.
The
FileIOintegration is cool, but in my experience, the old schoolreadtablefromDataFramesis a much more reliable CSV parser thanTextParseorCSV. For example, neitherTextParsenorCSVare able to read the actual data I am working with (I've included just a subset of the columns in the MWE above), whilereadableworks fine, aside from the type issues.Ah, interesting. Can you open issues about those problems, maybe even in TextParse.jl and CSV.jl? I think the maintainers of DataFrames.jl and DataTables.jl plan to do away with
readtable, so it would be quite crucial that the replacement parsers are able to handle the kind of data you are using.Alright, the . lifting doesn't work because some of the methods in base are not type stable. It might be enough to add a return type annotation here, but not sure.
queryverse/DataValues.jl#15 has a new method for the
Dateconstructor, once that is merged and tagged, the non-dot-broadcasting version should just work.Yes, I have opened issues for CSV parsing already:
JuliaData/CSV.jl#86
queryverse/TextParse.jl#19here is a related issue, perhaps with the handling of the
splitfunction onDataValues:julia> df = DataFrame(id = [1, 1, 2, 2], val = @data(["a,b", NA, "c,d", "e,f"])) julia> @from i in df begin @where !isnull(i.val) @select {i.id, names = split(i.val, ",")} into i @from j in i.names @select {i.id, newval = j} @collect DataFrame end ERROR: type UnionAll has no field parameters Stacktrace: [1] select_many(::Query.EnumerableSelect{NamedTuples._NT_id_names{DataValues.DataValue{Int64},_} where _,Query.EnumerableWhere{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Query.EnumerableIterable{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},IterableTables.DataFrameIterator{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Tuple{DataArrays.DataArray{Int64,1},DataArrays.DataArray{String,1}}}},##49#53},##50#54}, ::##51#55, ::Expr, ::Function, ::Expr) at /Users/tcovert/.julia/v0.6/Query/src/enumerable/enumerable_selectmany.jl:51 julia> @from i in df begin @where !isnull(i.val) @select {i.id, names = split.(i.val, ",")} into i @from j in i.names @select {i.id, newval = j} @collect DataFrame end ERROR: Stacktrace: [1] query(::DataValues.DataValue{Array{SubString{String},1}}) at /Users/tcovert/.julia/v0.6/Query/src/sources/source_iterable.jl:6 [2] start(::Query.EnumerableSelectMany{NamedTuples._NT_id_newval{DataValues.DataValue{Int64},Array{SubString{String},1}},Query.EnumerableSelect{NamedTuples._NT_id_names{DataValues.DataValue{Int64},DataValues.DataValue{Array{SubString{String},1}}},Query.EnumerableWhere{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Query.EnumerableIterable{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},IterableTables.DataFrameIterator{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Tuple{DataArrays.DataArray{Int64,1},DataArrays.DataArray{String,1}}}},##57#62},##58#63},##60#65,##61#66}) at /Users/tcovert/.julia/v0.6/Query/src/enumerable/enumerable_selectmany.jl:67 [3] macro expansion at /Users/tcovert/.julia/v0.6/IterableTables/src/integrations/dataframes.jl:91 [inlined] [4] _filldf(::Tuple{DataArrays.DataArray{Int64,1},Array{Array{SubString{String},1},1}}, ::Query.EnumerableSelectMany{NamedTuples._NT_id_newval{DataValues.DataValue{Int64},Array{SubString{String},1}},Query.EnumerableSelect{NamedTuples._NT_id_names{DataValues.DataValue{Int64},DataValues.DataValue{Array{SubString{String},1}}},Query.EnumerableWhere{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Query.EnumerableIterable{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},IterableTables.DataFrameIterator{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Tuple{DataArrays.DataArray{Int64,1},DataArrays.DataArray{String,1}}}},##57#62},##58#63},##60#65,##61#66}) at /Users/tcovert/.julia/v0.6/IterableTables/src/integrations/dataframes.jl:79 [5] _DataFrame(::Query.EnumerableSelectMany{NamedTuples._NT_id_newval{DataValues.DataValue{Int64},Array{SubString{String},1}},Query.EnumerableSelect{NamedTuples._NT_id_names{DataValues.DataValue{Int64},DataValues.DataValue{Array{SubString{String},1}}},Query.EnumerableWhere{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Query.EnumerableIterable{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},IterableTables.DataFrameIterator{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Tuple{DataArrays.DataArray{Int64,1},DataArrays.DataArray{String,1}}}},##57#62},##58#63},##60#65,##61#66}) at /Users/tcovert/.julia/v0.6/IterableTables/src/integrations/dataframes.jl:119 [6] DataFrames.DataFrame(::Query.EnumerableSelectMany{NamedTuples._NT_id_newval{DataValues.DataValue{Int64},Array{SubString{String},1}},Query.EnumerableSelect{NamedTuples._NT_id_names{DataValues.DataValue{Int64},DataValues.DataValue{Array{SubString{String},1}}},Query.EnumerableWhere{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Query.EnumerableIterable{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},IterableTables.DataFrameIterator{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Tuple{DataArrays.DataArray{Int64,1},DataArrays.DataArray{String,1}}}},##57#62},##58#63},##60#65,##61#66}) at /Users/tcovert/.julia/v0.6/IterableTables/src/integrations/dataframes.jl:128 [7] collect(::Query.EnumerableSelectMany{NamedTuples._NT_id_newval{DataValues.DataValue{Int64},Array{SubString{String},1}},Query.EnumerableSelect{NamedTuples._NT_id_names{DataValues.DataValue{Int64},DataValues.DataValue{Array{SubString{String},1}}},Query.EnumerableWhere{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Query.EnumerableIterable{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},IterableTables.DataFrameIterator{NamedTuples._NT_id_val{DataValues.DataValue{Int64},DataValues.DataValue{String}},Tuple{DataArrays.DataArray{Int64,1},DataArrays.DataArray{String,1}}}},##57#62},##58#63},##60#65,##61#66}, ::Type{DataFrames.DataFrame}) at /Users/tcovert/.julia/v0.6/Query/src/sinks/sink_type.jl:2as you might guess, this is related to the question I posed over at
IndexedTableshere: JuliaData/IndexedTables.jl#70edit: apparently this works though:
julia> @from i in df begin @where !isnull(i.val) @select {i.id, names = split.(get(i.val), ",")} into i @from j in i.names @select {i.id, newval = j} @collect DataFrame end 6×2 DataFrames.DataFrame │ Row │ id │ newval │ ├─────┼────┼────────┤ │ 1 │ 1 │ "a" │ │ 2 │ 1 │ "b" │ │ 3 │ 2 │ "c" │ │ 4 │ 2 │ "d" │ │ 5 │ 2 │ "e" │ │ 6 │ 2 │ "f" │why does this query need both array broadcasting and a
getcall to work?edit 2: ok now I see that this works as well:
julia> @from i in df begin @where !isnull(i.val) @select {i.id, names = split(get(i.val), ",")} into i @from j in i.names @select {i.id, newval = j} @collect DataFrame end 6×2 DataFrames.DataFrame │ Row │ id │ newval │ ├─────┼────┼────────┤ │ 1 │ 1 │ "a" │ │ 2 │ 1 │ "b" │ │ 3 │ 2 │ "c" │ │ 4 │ 2 │ "d" │ │ 5 │ 2 │ "e" │ │ 6 │ 2 │ "f" │so to answer my own question, no array broadcasting is needed, just
get@tcovert I can close this, right?
I have one idea for the
splitstory here queryverse/DataValues.jl#34. Probably a really bad idea, if you have an opinion, I'd be interested to hear it over there.If I missed some open issue here, just let me know and I'll reopen.
Suppose I have a DataFrame with two fields:
idxanddate. Thedatefield has missing values (in theDataFramessense) and is currently stored in the DataFrame as a string. Is there a query statement that I can write which parses the string into a date? I tried something like this:but got an error like this:
I also tried a version with no dot-broadcasting:
and got this error:
is what I am trying to do possible? if so, what am I doing wrong?
thanks in advance for any suggestions you can offer.
here is some example data to apply the code to above: https://www.dropbox.com/s/kgiicawhegmtavc/query_example.csv?dl=0