为什么这些按位正则表达式在golang中匹配不同？

Given this binary data and two regular expressions, why does golang match them differently?

var (
    data = []byte{0x03, 0x00, 0x00, 0x88, 0x02, 0xf0, 0x80, 0x72, 0x03, 0x00, 0x79, 0x20, 0xdd, 0x39, 0x22, 0x4d, 0xcf, 0x6d, 0x17, 0x29, 0x02, 0xee, 0xe3, 0x7f, 0x5e, 0xca, 0x17, 0x62, 0xc8, 0x56, 0x24, 0x01, 0x1e, 0x9f, 0xa0, 0x96, 0xc6, 0x4f, 0xbb, 0xa2, 0x51, 0x7b, 0xbf, 0x33, 0x31, 0x00, 0x00, 0x05, 0x4c, 0x00, 0x00, 0x1a, 0x6a, 0x00, 0x00, 0x03, 0x8d, 0x34, 0x00, 0x00, 0x00, 0x00, 0x04, 0x14, 0x81, 0xe7, 0x86, 0xa7, 0x76, 0x51, 0x02, 0x9d, 0x18, 0x09, 0xff, 0xde, 0xde, 0x05, 0x51, 0x02, 0x9d, 0x18, 0x0a, 0x89, 0xee, 0xf8, 0x81, 0x53, 0x51, 0x02, 0x9d, 0x18, 0x0b, 0x82, 0xce, 0xef, 0xad, 0x63, 0x51, 0x02, 0x9d, 0x18, 0x0c, 0x00, 0x00, 0x04, 0xe8, 0x89, 0x69, 0x00, 0x12, 0x00, 0x00, 0x00, 0x00, 0x89, 0x6a, 0x00, 0x13, 0x00, 0x89, 0x6b, 0x00, 0x04, 0x00, 0x00, 0xb4, 0x62, 0x00, 0x00, 0x00, 0x00, 0x72, 0x03, 0x00, 0x00}

    regexA = regexp.MustCompile(`\x72\x03(?P<TLen>[\x00-\xFF]{2})(?P<Payload>[\x00-\xFF]+)`)
    regexB = regexp.MustCompile(`\x72\x03(?P<TLen>[\x00-\xFF]{2})(?P<Payload>.+)`)
)

Shouldn't regexA's [\x00-\xFF]+ match as good as regexB's .+?

I'm testing using this:

func main() {
    log.Printf("testing regexA")
    applyRegex(regexA)
    log.Printf("testing regexB")
    applyRegex(regexB)
}

func applyRegex(r *regexp.Regexp) {

    matches := r.FindAllSubmatch(data, -1)
    groups := r.SubexpNames()

    for mIdx, match := range matches {
        findings := smap{}
        for idx, submatch := range match {
            findings[groups[idx]] = fmt.Sprintf("% x", submatch)
        }
        log.Printf("match #%d: %+v", mIdx, findings)
    }
}

And getting this output:

2009/11/10 23:00:00 testing regexA
2009/11/10 23:00:00 match #0: map[:72 03 00 79 20 TLen:00 79 Payload:20]
2009/11/10 23:00:00 testing regexB
2009/11/10 23:00:00 match #0: map[:72 03 00 79 20 dd 39 22 4d cf 6d 17 29 02 ee e3 7f 5e ca 17 62 c8 56 24 01 1e 9f a0 96 c6 4f bb a2 51 7b bf 33 31 00 00 05 4c 00 00 1a 6a 00 00 03 8d 34 00 00 00 00 04 14 81 e7 86 a7 76 51 02 9d 18 09 ff de de 05 51 02 9d 18 TLen:00 79 Payload:20 dd 39 22 4d cf 6d 17 29 02 ee e3 7f 5e ca 17 62 c8 56 24 01 1e 9f a0 96 c6 4f bb a2 51 7b bf 33 31 00 00 05 4c 00 00 1a 6a 00 00 03 8d 34 00 00 00 00 04 14 81 e7 86 a7 76 51 02 9d 18 09 ff de de 05 51 02 9d 18]

Is parsing binary data using regex not an actual possibility?

Playground link: https://play.golang.org/p/gHWqeyPuPNJ

展开全部

写回答
好问题 0 提建议
关注问题
分享
邀请回答
编辑收藏删除结题
收藏举报

1条回答默认最新

关注

码龄粉丝数原力等级 --

被采纳

被点赞

采纳率
douxian3170 2019-01-29 07:58
关注
The regexp documentation states that:

All characters are UTF-8-encoded code points.

So I think the character class ranges will be unicode code points ranges and not byte ranges. Also the matched text is treated as UTF-8 so i'm not sure what will happen at data = []byte{... , 0xdd, ...}. It might be decoded as a code point greater than 0xff. So i'm not sure how well it will work to use the standard library regexp package to do binary matching.

A side note (?P<Payload>.+) will not match all code points but (?P<Payload>(?s).+) will. s flag is:

let . match (default false)

Hope that helps

本回答被题主选为最佳回答 , 对您是否有帮助呢?

解决无用
评论打赏
分享
举报
编辑

预览
轻敲空格完成输入
显示为

卡片

标题

链接
评论

按下Enter换行，Ctrl+Enter发表内容

编辑

预览

报告相同问题？

关注问题

Golang 正则表达式
2024-11-01 14:39

MeteionY的博客如果没有匹配的字符串，那么它回...ReplaceAllStringFunc 返回一个字符串的副本，其中正则表达式的所有匹配项都已替换为指定函数的返回值。FindStringSubmatch 返回包含匹配项的字符串切片，包括来自捕获组的字符串。
go 正则表达式分组匹配,如何在Golang正则表达式中捕获组功能？
2021-01-17 09:20

臭人鹏的博客 I'm porting a library from Ruby to Go, and have just discovered that regular expressions in Ruby are not compatible with Go (google RE2). It's come to my attention that Ruby & Java (plus other lan...
【正则表达式】golang 正则表达式的正确使用姿势
2022-11-21 01:51

自驱的博客【代码】【正则表达式】golang 正则表达式的正确使用姿势。
go 正则表达式分组匹配_基础知识 - Golang 中的正则表达式
2021-02-27 05:26

weixin_39794347的博客 正则表达式里的分枝条件指的是有几种规则，如果满足其中任意一种规则都应该当成匹配，具体方法是用|把不同的规则分隔开。听不明白？没关系，看例子： 0\d{2}-\d{8}|0\d{3}-\d{7}这个表达式能匹配两种以连字号分隔的...
golang-正则表达式
2023-06-19 05:48

雨师@的博客 Regexp包提供了16个方法，用于匹配正则表达式搜索结果。方法名满足如下正则表达式： Find(All)?(String)?(Submatch)?(Index)?Regexp包提供了16个方法，用于匹配正则表达式搜索结果。方法名满足如下正则表达式： ...
浅析golang 正则表达式
2020-10-14 09:00

正则表达式是用于匹配字符串中字符组合的模式，它是文本处理中非常强大的工具。在Go语言中，正则表达式得到了良好的支持，并且使用起来相对直观。Go的正则表达式库是`regexp`包，它提供了正则表达式的支持，可以从...
golang 正则中括号匹配_TechRepo | 正则表达式
2021-01-18 21:46

山有灬扶苏的博客 TechRepo是由软件学院学生科协推出的技术分享系列推送，每周会进行一次更新，旨在为同学们提供一个...Regular Expression正/则/表/达/式基本语法介绍No.0序言在搜索中，我们都用过正则表达式。如?通配符匹配文件名中...
go 正则表达式分组匹配_golang正则表达式regexp示例大全
2021-01-17 09:20

萧竹声的博客 ————————————————————// 判断在 b 中能否找到正则表达式 pattern 所匹配的子串// pattern：要查找的正则表达式// b：要在其中进行查找的 []byte// matched：返回是否找到匹配项// err：返回查找...
golang使用正则表达式解析网页
2020-10-24 04:44

在本文中，将介绍如何使用Go语言（通常称为Golang）编写程序以使用正则表达式解析网页。Go语言以其简洁的语法和高效的性能被广泛应用于网络编程中。正则表达式是一种强大的文本处理工具，常用于匹配和操作字符串。 ...
非零基础自学Golang 第16章 正则表达式 16.1 正则表达式介绍 & 16.2 正则表达式语法
2022-12-22 06:37

Ding Jiaxiong的博客非零基础自学Golang 第16章 正则表达式 16.1 正则表达式介绍 & 16.2 正则表达式语法
没有解决我的问题, 去提问

为什么这些按位正则表达式在golang中匹配不同？

1条回答 默认 最新

1条回答默认最新