duanli9569 2018-08-07 12:48 采纳率: 100%
浏览 106
已采纳

解析UTF8文本时反斜杠导致问题

I used windows cmd dir /s command to get a list of all pdf files.
Now I want to parse the text and create a simple table that I can copy paste to Excel (after some more text parsing is done).
Just to explain why I don't want to do this in Excel, I need to use the levenshtein function to unifrom/group similar items. But that is not part of the question, I can do that later myself.

My first attempt was regex.

$re = '/(\d{4})\\(\d{2})\\(\d{2})\\(.+?)\\(\d+)-(.+?)\\(.+?) -/m';
$re = '/(\d{4}).(\d{2}).(\d{2}).(.+?).(\d+)-(.+?)\\\\(.+?) -/m';

Non of them works when I run them on 3v4l but on on regex101 the first one works and the second one a simplified version where dot replaces a backslash.
But unfortunatly I can't parse the last bit without a backslash.

My second attempt was simple explode on backslash But that didn't work

$arr = explode("
", $str);

foreach($arr as $line){
    $parts = explode('\\', $line);
    var_dump($parts);
}

https://3v4l.org/JZ8gR
Because the backslash is used as an escape (I think) in the string.
So I tried to replace the backslash with dash.

$arr = explode("
", str_replace("\\", "-", $str));
var_dump($arr);/*

https://3v4l.org/Xcs0G
But yet again my text finds a way to beat me.

Full text can be found in any of the links above. A smaler example:

H:\Dokument\Avvikelser\2018\08\03\ALIMENTOS DEL MEDITERRANE\243715000-Vattenmelon\Kvalitets fel - avvikelse27210.pdf
H:\Dokument\Avvikelser\2018\08\06\GRÖNSAKSMÄSTARNA SVERIGE\000233003-Kålrötter 6kg RB\Kvalitets fel - avvikelse27245.pdf
H:\Dokument\Avvikelser\2018\08\06\GRÖNSAKSMÄSTARNA SVERIGE\000223005-Isbergssall. påse RB\Kvalitets fel - avvikelse27244.pdf
H:\Dokument\Avvikelser\2018\08\06\GRÖNSAKSMÄSTARNA SVERIGE\223005000-Isberg påse RB\Kvalitets fel - avvikelse27272.pdf
H:\Dokument\Avvikelser\2018\08\06\TERRA NATURA INTERNATIONA\277711000-Tomat kvist 5kg\ - avvikelse27270.pdf
H:\Dokument\Avvikelser\2018\08\06\TERRA NATURA INTERNATIONA\277711000-Tomat kvist 5kg\Kvalitets fel - avvikelse27270.pdf
H:\Dokument\Avvikelser\2018\08\06\LCT i Skåne\221715000-Ingefära 5kg\Kvalitets fel - avvikelse27279.pdf

What I expect is a each line parsed in such a way that the backslash is not causing issues.
Example:

["H:", "Dokument", "Avvikelser",", "2018", "08", "06", "LCT i Skåne", "221715000", "Ingefära 5kg", "Kvalitets fel", "avvikelse27279.pdf"]

But as the regex implies, I don't need all parts of the string.

["2018", "08", "06", "LCT i Skåne", "221715000", "Ingefära 5kg", "Kvalitets fel"]

is enough.

EDIT: I'm fine with using EOD or " or any other way to initiate the string. But since a ' is used in the text that can't be used.

  • 写回答

1条回答 默认 最新

  • douguai7291 2018-08-07 13:20
    关注

    Use Nowdoc like this, surrounding the "END WORD" with single quotes :

    $str = <<<'EOD'
    H:\Dokument\Avvikelser\2018\08\03\ALIMENTOS DEL MEDITERRANE\243715000-Vattenmelon\Kvalitets fel - avvikelse27210.pdf
    H:\Dokument\Avvikelser\2018\08\06\GRÖNSAKSMÄSTARNA SVERIGE\000233003-Kålrötter 6kg RB\Kvalitets fel - avvikelse27245.pdf
    H:\Dokument\Avvikelser\2018\08\06\GRÖNSAKSMÄSTARNA SVERIGE\000223005-Isbergssall. påse RB\Kvalitets fel - avvikelse27244.pdf
    H:\Dokument\Avvikelser\2018\08\06\GRÖNSAKSMÄSTARNA SVERIGE\223005000-Isberg påse RB\Kvalitets fel - avvikelse27272.pdf
    H:\Dokument\Avvikelser\2018\08\06\TERRA NATURA INTERNATIONA\277711000-Tomat kvist 5kg\ - avvikelse27270.pdf
    H:\Dokument\Avvikelser\2018\08\06\TERRA NATURA INTERNATIONA\277711000-Tomat kvist 5kg\Kvalitets fel - avvikelse27270.pdf
    H:\Dokument\Avvikelser\2018\08\06\LCT i Skåne\221715000-Ingefära 5kg\Kvalitets fel - avvikelse27279.pdf
    EOD;
    
    $re = '/(\d{4})\\\\(\d{2})\\\\(\d{2})\\\\(.+?)\\\\(\d+)-(.+?)\\\\(.+?) -/m';
    $res = preg_match($re, $str, $m);
    
    print_r($m);
    
    本回答被题主选为最佳回答 , 对您是否有帮助呢?
    评论

报告相同问题?

悬赏问题

  • ¥15 关于#Java#的问题,如何解决?
  • ¥15 加热介质是液体,换热器壳侧导热系数和总的导热系数怎么算
  • ¥15 想问一下树莓派接上显示屏后出现如图所示画面,是什么问题导致的
  • ¥100 嵌入式系统基于PIC16F882和热敏电阻的数字温度计
  • ¥15 cmd cl 0x000007b
  • ¥20 BAPI_PR_CHANGE how to add account assignment information for service line
  • ¥500 火焰左右视图、视差(基于双目相机)
  • ¥100 set_link_state
  • ¥15 虚幻5 UE美术毛发渲染
  • ¥15 CVRP 图论 物流运输优化