duanpin9531 2018-08-22 18:53
浏览 50
已采纳

Go中最快的数据加载方式

I need to regularly load over 300'000 rows x 78 columns of data into my Go program.

Currently I use (import github.com/360EntSecGroup-Skylar/excelize):

xlsx, err := excelize.OpenFile("/media/test snaps.xlsm")
if err != nil {
    fmt.Println(err)
    return
}

//read all rows into df
df := xlsx.GetRows("data")

It takes about 4 minutes on a decent PC using Samsung 960 EVO Series - M.2 Internal SSD.

Is there a faster way to load this data? Currently it takes me more time to read the data than processing it. I'm also opened to other file formats.

  • 写回答

1条回答 默认 最新

  • doupi7619 2018-08-22 20:11
    关注

    As suggested in the comments, instead of using the XLS format, use a custom, fast data format for reading and writing your table.

    In the most basic case, just write the number of columns and rows to a binary file, then write all the data in one go. This will be very fast, I have created a little example here which just writes 300.000 by 40 float32s to a file and reads them back. On my machine this takes about 400ms and 250ms (notice that the file is hot in cache after writing it, may take longer on initial read).

    package main
    
    import (
        "encoding/binary"
        "os"
    
        "github.com/gonutz/tic"
    )
    
    func main() {
        const (
            rowCount = 300000
            colCount = 40
        )
        values := make([]float32, rowCount*colCount)
        func() {
            defer tic.Toc()("write")
            f, _ := os.Create("file")
            defer f.Close()
            binary.Write(f, binary.LittleEndian, int64(rowCount))
            binary.Write(f, binary.LittleEndian, int64(colCount))
            check(binary.Write(f, binary.LittleEndian, values))
        }()
        func() {
            defer tic.Toc()("read")
            f, _ := os.Open("file")
            defer f.Close()
            var rows, cols int64
            binary.Read(f, binary.LittleEndian, &rows)
            binary.Read(f, binary.LittleEndian, &cols)
            vals := make([]float32, rows*cols)
            check(binary.Read(f, binary.LittleEndian, vals))
        }()
    }
    
    func check(err error) {
        if err != nil {
            panic(err)
        }
    }
    
    本回答被题主选为最佳回答 , 对您是否有帮助呢?
    评论

报告相同问题?

悬赏问题

  • ¥15 求差集那个函数有问题,有无佬可以解决
  • ¥15 【提问】基于Invest的水源涵养
  • ¥20 微信网友居然可以通过vx号找到我绑的手机号
  • ¥15 寻一个支付宝扫码远程授权登录的软件助手app
  • ¥15 解riccati方程组
  • ¥15 display:none;样式在嵌套结构中的已设置了display样式的元素上不起作用?
  • ¥15 使用rabbitMQ 消息队列作为url源进行多线程爬取时,总有几个url没有处理的问题。
  • ¥15 Ubuntu在安装序列比对软件STAR时出现报错如何解决
  • ¥50 树莓派安卓APK系统签名
  • ¥65 汇编语言除法溢出问题