dousi8237
dousi8237
2010-09-30 15:18

如何编写preg_match_all只是为了获取一个特定的元素?

已采纳

Until the website give me an access to his API, i need to display only 2 things from this website :

What i want to grab // Example on a live page

Those 2 things are contained in a div :

<div style="float: right; margin: 10px;">
here what i want to display on my website
</div>

The problem is that i found an example on stackoverflow, but i never wrote preg_match before. How to do this with the data i want to grabb ? Thank you

<?php   $html = file_get_contents($st_player_cv->getUrlEsl());

preg_match_all(
    'What do i need to write here ?',
    $html,
    $posts, // will contain the data
    PREG_SET_ORDER // formats data into an array of posts
);

foreach ($posts as $post) {
    $premium = $post[1];
    $level = $post[2];

    // do something with data
}
  • 点赞
  • 写回答
  • 关注问题
  • 收藏
  • 复制链接分享
  • 邀请回答

3条回答

  • duanchao1002 duanchao1002 11年前

    The DOM way to do it would be

    libxml_use_internal_errors(TRUE);
    $dom = new DOMDocument;
    $dom->loadHTMLFile('http://www.esl.eu/fr/player/5178309/');
    libxml_clear_errors();
    
    $xPath = new DOMXPath($dom);
    $nodes = $xPath->query('//div[@style="float: right; margin: 10px;"]');
    foreach($nodes as $node) {
        echo $node->nodeValue, PHP_EOL;
    }
    

    but there is a whole slew of JavaScript in the page that modifies the DOM heavily after the page was loaded. Since any PHP script based fetching will not execute any JavaScript, the style we search for in the XPath does not exist yet and we won't get any results (the Regex suggesed by Hannes doesn't work for the same reason). Neither do the level numbers on the badge exist yet.

    As Wrikken pointed out in the comments, there also seems to be some mechanism to block certain requests. I had the message once, but I am not sure what triggers it, because I could also fetch page on several occasions.

    To cut a long story short: you cannot achieve what you are trying to do with this page.

    点赞 评论 复制链接分享
  • dongpuchao1680 dongpuchao1680 11年前

    If you want something more generic

      preg_match('/<div[^>]+?>(.*?)<\/div>/', $myhtml, $result);
      echo $result[1] . "
    ";
    

    $myhtml contains the code html you have to analyze. $result is the array that contains the regexp and () content after the regular expression was applied. $result[1] will give you what is between the <div ... > and </div>.

    This way, even if the <div differs (class name change or different attributes), it'll still work.

    点赞 评论 复制链接分享
  • dongtao1262 dongtao1262 11年前

    this regex '#<div style="float: right; margin: 10px;">(.*)</div>#' should do the trick (yeah) but i would advice you to use DOM & XPath.

    edit:

    Here is an Xpath / DOM Example:

    $html = <<<HTML
    <html>
    <body>
        <em>nonsense</em>
        <div style="float: right; margin: 10px;"> here what i want to display on my website </div>
        <div> even more nonsense </div>
    </body>
    </html>
    
    HTML;
    
    $doc = new DOMDocument();
    $doc->loadHTML($html);
    $xpath = new DOMXpath($doc);
    $elements = $xpath->query('//div[@style="float: right; margin: 10px;"]');
    echo $elements->item(0)->nodeValue;
    
    点赞 评论 复制链接分享

相关推荐